Building or Buying? AI for the Scholarly Ecosystem

Feb 14, 2024 · 0 min read
Abstract
The title of this panel sounds like you have a choice between options but, in fact, most academic projects that employ machine learning or artificial intelligence require both buying and building. Representing the ‘building’ side of the panel, I survey real-world AI/ML projects at an academic research library: automating the segmentation of television news broadcasts, using retrieval-augmented generation to produce credible Wikipedia articles about women religious leaders, experimenting with diffusion models for a course on sustainable fashion, and training a supervised model to distinguish the anchors, reporters, and experts who contribute to news programming. Drawing on these examples, I review the platforms we return to again and again—Databricks, Google Colab, AWS machine learning services, Hugging Face, vector databases, and Labelbox—and the mostly open-source protocols (packages, libraries, and APIs such as Top2Vec, the OpenAI API, LangChain, Marvin, Unstructured, and MLflow) that hold these projects together. I conclude that the calculus should still be weighted toward buying: the field moves so quickly that home-built applications risk constant recoding as APIs change—or being superseded outright by the market.
Date
Feb 14, 2024 12:00 AM — 1:00 AM
Event
NISO Plus 2024
Location

Baltimore, MD

events
Clifford B. Anderson
Authors
Director of the Divinity Library
My research interests include the study of algorithms as cultural artifacts, computational thinking in the humanities, large-scale textual analysis of narrative data, and the religious dimensions of intellectual property.