The Smithsonian Is Using AI to Find the People History Forgot

By Ellis Ward, Resident Expert, AI Systems and Research Norms, RFTUV


A silversmith places a newspaper ad in 1774. A creamer matching his described work sits in a museum case 250 years later. Until recently, connecting those two facts required a researcher who already knew to look, knew where both catalogs lived, and had enough hours to manually cross-reference them. Most connections like this were never made. The Revolution Crossroads project, a collaboration between the Smithsonian Institution and the Library of Congress, is applying AI models to close that gap at scale, in time for the nation's 250th anniversary in 2026.


This piece covers the method, the data architecture, the human review layer, and the honest limits the project's own team acknowledges.


What the Project Is Actually Doing


Revolution Crossroads began applying AI models in spring 2026, according to reporting by AP. The goal, as described by the institutions, is to let the public and scholars search Smithsonian collections and draw connections across both institutions to understand how communities formed and functioned at the nation's founding.


The project is centered on roughly 10,000 objects linked to people who lived between 1770 and 1810. The materials include newspapers, letters, military records, artifacts, maps, and portraits. The problem is not that these materials are secret. They are separated across institutions with different catalog systems, different metadata standards, and no common index.


The AI layer creates a data bridge. Pattern recognition runs across documents and objects that would take, as Kenneth Cohen told Smithsonian Magazine in a July 30, 2026 piece by Kelly Revak, a historian a lifetime to cross-reference manually.


The two institutions bring different collections to the table. Cohen noted that the Library of Congress holdings are mostly documentary: written records, drawn images. The Smithsonian's holdings are mostly three-dimensional: objects, artifacts, material culture. Putting them in conversation, Cohen said, reveals patterns that neither institution's collection shows alone.


The Data Architecture


The project did not simply run existing catalog data through a model. Team members had to digitize and redigitize objects specifically for AI processing. That matters methodologically: the quality and consistency of input data shapes what the model can and cannot surface.


The dataset is open. It lives on Hugging Face at huggingface.co/RevolutionCrossroads and includes Smithsonian records, Library of Congress records, and materials from other institutions, along with metadata and media suited for digital humanities, machine learning, and natural language processing work. Becky Kobberod, chief digital and innovation officer at the Smithsonian, confirmed the datasets are freely available to outside researchers.


The public-facing search tool is at si.edu. The Smithsonian, founded in 1846, now comprises 21 museums and galleries including the National Zoo. The Library of Congress documented the project's development in its Signal blog in September 2026, framing the 2026 semiquincentennial as the occasion for this kind of cross-institutional effort.


The Silversmith Example: What Connection Looks Like in Practice


The project's clearest illustration of method is the silversmith case. A craftsman ran newspaper ads around 1774 describing a creamer. A creamer similar to his description sits at the National Museum of American History. Previously, connecting the ad to the object required a manual search across separate catalogs: finding the newspaper, reading the ad, recognizing the object description, then knowing to look in the museum's holdings.


Revolution Crossroads surfaces that link automatically.


Natalie Buda Smith, director of digital strategy at the Library of Congress, described the creamer as illustrating that there's a richness out there and you might find part of the story at the Smithsonian and part of the story at the library. She added: Open up those doors, because they found it in the Smithsonian doesn't mean it's the end of the story.


That framing matters. The system is not producing definitive attributions. It is producing leads. The distinction between a match and a conclusion belongs to the historian.


Who Gets Surfaced, and Why That Is the Point


The project's explicit aim is to recover names and lives that do not appear in the conventional historical record. Kobberod described an ideal scenario where the project is actually adding to the public record. She elaborated: Maybe they're not Thomas Jefferson or George Washington, but they did live in a community and build that community, and that name is now known.


Cohen put the research framing directly: There are casts of characters that still have to be uncovered, and AI allows us to find references and patterns across huge amounts in ways that would take a historian a lifetime.


The Library of Congress's theme for its 2026 semiquincentennial work is It's Your Story, signaling that the institutional intent is explicitly inclusive. The Smithsonian's Secretary Lonnie G. Bunch III described the institution as the place the public goes to understand itself through history, culture, art, and science using the best available technologies.


For researchers in early American history, genealogy, and social history, this matters concretely. Communities that did not generate the kind of documentary record that survives in letters and diaries may appear in material culture, in newspaper ads, in military records, or in objects that the AI can now link across institutions.


The Human Review Layer: Why Historians Are Not Optional


Revolution Crossroads is not using current frontier models. That is a deliberate choice, and Kobberod explained the reasoning in terms of trust and verification.


Historians review the model's outputs. Their role is to catch nuance the model cannot hold. Kobberod gave two examples of the kind of correction that comes back from historians: actually that is a known nickname for so-and-so, and that is a generic term that's used for X, Y, Z. She stated plainly: That is the nuance that a historian is going to know that a machine is not.


Cohen echoed the limit from the data side: AI can't counter the biases and perspectives that are embedded in the evidence, so we've got to bring that human awareness. The historical record of 1770 to 1810 reflects who was documented and by whom. An AI model trained on that record inherits those gaps. It can surface patterns within what survives; it cannot reconstruct what was never recorded.


Kobberod described her team's orientation as a healthy fear of AI, a phrase worth taking seriously. The project is not treating the technology as neutral infrastructure. She said: It's not going away. We do believe we need to lean in and figure out the part that this technology plays in our work.


The project's transparency commitment is also notable. Kobberod said the team will be showing some of our experiments, and showing the outputs, and also showing when it's not working. That is an unusual public commitment for an institution-scale AI deployment, and it sets a useful benchmark.


Key Numbers in Context


Objects centered on 1770 to 1810 lives: about 10,000. Smithsonian museums and galleries: 21. Smithsonian founding year: 1846. AI models in use: not current frontier models. Project start, AI application phase: spring 2026. Public search tool: si.edu. Open dataset: huggingface.co/RevolutionCrossroads.


Limits and Counterpoints


The gap is in what survived, not only in how we search it. AI can surface connections within the existing record. It cannot recover what was never documented. Communities with low documentary survival may appear in fragments. Pattern recognition will find those fragments faster. It will not fill gaps that do not exist in the data.


Match is not attribution. The silversmith example illustrates a likely connection between an ad and an object. Establishing that connection as historical fact requires the kind of source criticism and contextual knowledge that historians do professionally. The model produces candidates, not conclusions.


Scaling requires ongoing curation. The digitization and redigitization work the team describes is labor-intensive. The open dataset on Hugging Face invites outside researchers to extend the work, which distributes the burden. But the quality of that extension depends on metadata standards that not all contributing institutions maintain consistently.


Practical Takeaway


If you are a researcher, educator, or engaged reader with interest in the 1770 to 1810 period, search at si.edu. The public interface is live and designed for general use, not only specialists. Check the Hugging Face repository. The open dataset includes metadata and media suited for independent digital humanities work. Read the Library of Congress Signal post for the methodological context on how the data bridge between institutions was constructed. Watch the Can AI Help Historians video linked from the Smithsonian Magazine piece for a plain-language account of what the system is and is not doing.


The project's framing is accurate to the technology: AI as a pattern-finding layer, historians as the interpretive layer, and the public as the beneficiary of connections that would otherwise remain buried in separate catalogs. The 250th anniversary is a deadline, but the infrastructure being built here is designed to outlast it.


Sources and further reading


AP via Phys.org (2026): Smithsonian AI and American Revolution artifacts: https://phys.org/news/2026-10-smithsonian-ai-artifacts-american-revolution.html


Library of Congress, The Signal (September 2026): Revolution Crossroads: https://blogs.loc.gov/thesignal/2026/09/revolution-crossroads


Smithsonian Magazine, Office of Digital Innovation, Kelly Revak (July 30, 2026): Can AI Help Historians Discover the Untold Stories of the American Revolution: https://www.smithsonianmag.com/blogs/office-digital-innovation/2026/07/30/can-ai-help-historians-discover-the-untold-stories-of-the-american-revolution/


Revolution Crossroads open dataset, Hugging Face: https://huggingface.co/RevolutionCrossroads

Ellis Ward

Ellis Ward is Reporting from the Uncanny Valley's resident expert on AI systems and research norms. He explains how models are tested, where agents break down, and what a study does and does not prove. He lives in Albany.

Previous
Previous

Dutchess Pulls Its Flock Cameras After Residents Push Back

Next
Next

Sam Altman Says Treating AI Like a Religion Is a Safety Issue