AI makes it easier to explore fish files
Data scientists at Carolina speed the transformation of a state museum’s handwritten specimen notes into searchable data.
Around the corner from the North Carolina State Fairgrounds, tucked behind a gated entrance, sits an unassuming building with something incredible inside: 1.5 million fish specimens.
Welcome to the North Carolina Museum of Natural Sciences’ Ichthyology Collection. (Try saying it out loud. It’s pronounced ick-thee-ol-uh-jee.)

Lily Hughes, research curator of ichthyology, examines hogchoker in the ichthyology lab within the NCMNS Research Lab. (North Carolina Museum of Natural Sciences)
“Ichthyology is the study of fishes, which means all of the biodiversity of fishes,” says Lily Hughes, research curator for the collection. “There’s about 37,000 species of fish that live in our planet today.”
Inside the collections room — sometimes called “the range” by staff — large stainless steel tanks line the walls. Narrow hallways cut through rows and rows of metal shelves engulfed with jars of sea life dating back to 1850. Eels, pufferfish, sturgeon, even a 13-foot thresher shark. It’s all there, just waiting to be explored.
“We are a lending institution,” Hughes explains. “We lend specimens out to researchers so they can ask their own questions and learn more about all of the life on our planet.”
On the other side of the building, a carpeted room hosts even more shelves, but these hold something arguably more important than the specimens themselves: field notes.
Each fish in the collection has a corresponding set of handwritten data: who collected it, where and when it was captured and other information that gets uploaded into a public database. That’s how a scientist in Japan studying lamprey, for example, could find information needed for a paper they’re writing.

Gabriela Hogue, collections manager of ichthyology, poses for a portrait in the ichthyology collection within the NCMNS Research Lab. (North Carolina Museum of Natural Sciences)
“If we don’t have that information, then we just have a specimen without any data — and we can’t really tell you much about it,” says collections manager Gabriela Hogue. “It can be used for teaching. It can be used for dissecting or making models. But it can’t be used for deeper research or for conservation and management.”
Hogue and the staff spend hours uploading handwritten specimen notes into electronic databases. This process is the same for museums around the world, from NCMNS to the Natural History Museum in London.
The Renaissance Computing Institute at UNC-Chapel Hill is working to reduce the time it takes to do this.
RENCI “specializes in data science and cyber infrastructure, which essentially just means that we use different computational approaches to help other people do their research,” says Chris Bizon, director of analytics and data science at the institute.

Northern hogsucker (Hypentelium nigricans, NCSM 28784) photographed at the NCMNS Research Lab. (North Carolina Museum of Natural Sciences)
Bizon and his team are trying to develop an agent using a multimodal large language model — AI that can process multiple types of data like text, images, audio and video at once — to transcribe these handwritten notes and infer the latitude and longitude for where each specimen was discovered.
To build something big, Bizon starts small, spending time with Hogue and Hughes to uncover their concerns, their pain points and how to achieve accuracy within their transcription process. Most recently, RENCI brought a bunch of their teammates together to host a hackathon, where they spent a half-day coding the agent, assessing problems and sharing ideas for how to fix them.
“A lot of it is really around trying something, looking at what comes out, going back, modifying the approach,” Bizon says. “It’s just like any other project where we have some goal we’re trying to achieve and some different approaches we can take.”
A tool like this could be revolutionary for museums worldwide.
“It is going to be an absolute life-changing thing for collection managers everywhere,” Hogue says. “By minimizing that time, we can maximize our effort handling the specimens and globally providing the data so that researchers, educators, conservation managers, teachers all over the world can utilize [it].”







