Making more than 1 million historical NLM records easier to find
An AI-enabled workflow transformed a complex, manual migration into a scalable and cost-effective process.
The National Library of Medicine’s (NLM) User Services and Collections Division (USCD) had a valuable resource hiding in plain sight: more than a million dissertation and monograph records housed in IndexCat™, a legacy database. USCD sought to merge this important database with its current catalog, Alma, to make these resources more accessible to the public. ICF helped develop an approach that made the migration practical, scalable, and cost-effective.
Challenge
Bringing the records housed in IndexCat™ to Alma would improve the search experience for NLM users, but creating and maintaining those records manually would require an enormous investment of staff time. Such a large-scale modernization would be difficult to sustain alongside library staff’s day-to-day responsibilities. NLM turned to the USCD GenAI Accelerator program and ICF to explore whether AI could help modernize historical records while maintaining the accuracy and oversight required for library collections.
- ICF Fathom
- Anthropic Claude
- AWS Bedrock
- Cloud
- Data Modernization
Solution
After a review of LLMs available on the market at the time and testing of the best-fit systems, we selected Anthropic’s Claude Opus 3 and Sonnet 3.5 models, as they demonstrated the complex reasoning required to accomplish this project’s challenging migration tasks.
Using the Claude models within AWS Bedrock, we combined automation scripts with targeted generative AI prompts to move, translate, and update records. In the GenAI Accelerator program, we developed a proof of concept that handled all three tasks concurrently, securely, and safely in a FedRAMP-compliant environment.
Because Anthropic partners with AWS, we were able to leverage AWS’s batching functionality. We made our instructions as clear and concise as possible, then used AWS to “batch” the data—waiting to process it until cloud traffic was lower. In exchange for a small delay, AWS processed the data at a significantly reduced cost.
Importantly, the approach keeps a human in the loop. Librarians double-check the records Claude creates, assign barcodes and call numbers to the physical items, and make the items discoverable in Alma. Although work remains, USCD’s librarians no longer have to spend their time creating full catalog records by hand.
Results
This AI-driven process cut the average time to process a record in half, from 30 minutes to an estimated 15. Migrating more than 1.1 million records with AI cost dramatically less than having staff complete the same work manually.
Much of the savings came from using Claude tools. Claude was the most cost-effective and highest-performing of the models we trialed for the project, and its integration with AWS enabled us to batch-process data at a significantly reduced cost.
NLM benefits from greater efficiency and lower costs, while users benefit from a greatly improved search experience: one interface for the library’s entire body of medical literature, rather than two.
Final thought
The migration of IndexCat™ monograph and dissertation records to Alma was completed in September of 2025, but we’re continuing to work with NLM on a variety of other AI projects in the USCD GenAI Accelerator program. The benefit of the Accelerator is that it allows NLM to experiment with AI in a structured way, investing major resources only in prototypes that will have the biggest impact on the client’s staff and work.
At ICF, we’re able to do this kind of targeted, outcomes-driven work because we combine deep expertise in technology with broad domain experience. This blend of expertise and domain experience allows us to help federal agencies develop the right tool to solve the right problems at the right cost.