Building or Buying? AI for the Scholarly Ecosystem
Feb 14, 2024
·
0 min read
Abstract
The title of this panel sounds like you have a choice between options but, in fact, most academic projects that employ machine learning or artificial intelligence require both buying and building. Representing the ‘building’ side of the panel, I survey real-world AI/ML projects at an academic research library: automating the segmentation of television news broadcasts, using retrieval-augmented generation to produce credible Wikipedia articles about women religious leaders, experimenting with diffusion models for a course on sustainable fashion, and training a supervised model to distinguish the anchors, reporters, and experts who contribute to news programming. Drawing on these examples, I review the platforms we return to again and again—Databricks, Google Colab, AWS machine learning services, Hugging Face, vector databases, and Labelbox—and the mostly open-source protocols (packages, libraries, and APIs such as Top2Vec, the OpenAI API, LangChain, Marvin, Unstructured, and MLflow) that hold these projects together. I conclude that the calculus should still be weighted toward buying: the field moves so quickly that home-built applications risk constant recoding as APIs change—or being superseded outright by the market.
Date
Feb 14, 2024 12:00 AM — 1:00 AM
Event
NISO Plus 2024
Location
Baltimore, MD
