Ancestors and Algorithms: AI for Genealogy
Stuck on a family history brick wall? It's time to add the most powerful tool to your genealogy toolkit: Artificial Intelligence. Welcome to Ancestors and Algorithms, the definitive guide to revolutionizing your family tree research with AI.
Forget the hype and confusion. This isn't just another podcast about AI; this is your hands-on, step-by-step masterclass using AI. Each week, host and researcher Brian demystifies the technology and shows you exactly how to apply AI tools to find ancestors, analyze records, and solve your toughest genealogy puzzles.
We explore the incredible promise of AI while navigating its perils with an honest, practical approach. Learn to use AI as your personal research assistant—not a replacement for your own critical thinking.
Join us to learn how to:
- Break through brick walls using AI-driven analysis and data correlation.
- Transcribe old, hard-to-read documents, letters, and census records in minutes.
- Use ChatGPT, Gemini, and other Generative AI to draft biographies, summarize findings, and organize your research.
- Analyze DNA matches and historical records to uncover hidden family connections.
- Master prompts that get you accurate results and avoid AI "hallucinations."
- Discover the latest AI tech and digital tools for genealogists before anyone else.
Whether you're a beginner genealogist or a seasoned family historian, if you're ready to upgrade your research skills, this podcast is for you. Hit Follow now and turn AI into your ultimate secret weapon for uncovering your ancestry.
Ancestors and Algorithms: AI for Genealogy
Ep. 51: The Jewish Shtetl No Map Remembers
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
A Jewish ancestor's hometown was spelled three different ways across three documents. A 1985 algorithm and three AI tools narrowed it down. This episode shows exactly how, and exactly where the method stops working.
This episode is for you if you are researching Eastern European Jewish ancestors, if a hometown name on an old document does not match any map you can find, or if you want to see how Claude, Gemini, and NotebookLM actually perform on real historical documents rather than a marketing demo.
You will learn:
- How the Daitch-Mokotoff Soundex system, built in 1985 specifically for Eastern European Jewish names, still powers the free JewishGen Communities Database today
- Why Ellis Island inspectors did not change immigrants' names, and where name changes actually happened instead
- How to use Gemini via Google AI Studio to transcribe messy historical handwriting, and how to read its uncertainty flags
- How to use Claude to weigh competing candidate towns against the evidence, and what it takes to turn a strong guess into a confirmed fact
- How to use NotebookLM to correlate multiple documents with citations, and why that matters when no single source closes the gap
FREQUENTLY ASKED QUESTIONS
What is Daitch-Mokotoff Soundex? Daitch-Mokotoff Soundex is a phonetic name-matching system built in 1985 by Gary Mokotoff and Randy Daitch specifically to handle the spelling variation common in Eastern European Jewish surnames and place names. It still powers JewishGen's free Communities Database today.
Did immigration officials at Ellis Island change immigrants' names? No. Ellis Island inspectors worked from passenger manifests already prepared by the shipping line in Europe and did not write names down for the first time. When a name changed, it typically happened later, at the hands of the immigrant or a naturalization clerk.
How accurate is Gemini at transcribing old handwritten documents? As of August 2026, independent testing on eighteenth and nineteenth century English-language handwriting has measured Gemini's character error rate at around one to two percent. Early Russian-language testing has been promising too, though not yet benchmarked as rigorously.
What is the JewishGen Communities Database? It is a free, Soundex-searchable database that lets researchers search for a place name phonetically rather than by exact spelling, then filter results by province and by historical Jewish population, useful when a name has been transliterated inconsistently across records.
Can AI tools confirm an ancestor's exact hometown? Not always, and this episode is honest about that. AI tools can narrow a list of real candidates and help weigh which one is best supported by the evidence, but confirming an exact match usually still requires a document that names the specific family in connection to that place.
What AI tools are featured in this episode? Claude, for comparing and reasoning across multiple transcribed documents. Gemini via Google AI Studio, for handwriting transcription. NotebookLM, for source-grounded correlation across every document gathered, each claim cited to its source.
For my Australian and UK listeners: this same wave of Eastern European Jewish emigration built major communities in Melbourne, Sydney, and London's East End. The same free Daitch-Mokotoff Soundex search works from anywhere in the world, and the National Archives of Australia and The National Archives in the UK are essential next steps once a place name is narrowed down.
An honest note on the outcome: this is a Brick Wall episode. The leading candidate town is well supported but not confirmed, and the episode says so plainly rather than dressing up a strong guess as solved.
For a deeper walkthrough of this exact technique, including additional advanced prompts and record-collection guidance for Eastern European Jewish research, the Companion Guide for this episode is available at ancestorsandai.com.
Connect with Ancestors and Algorithms:
📧 Email: ancestorsandai@gmail.com
🌐 Website: https://ancestorsandai.com/
📘 Facebook Group: Ancestors and Algorithms: AI for Genealogy - www.facebook.com/groups/ancestorsandalgorithms/
Golden Rule Reminder: AI is your research assistant, not your researcher.
Join our Facebook group to share your AI genealogy breakthroughs, ask questions, and connect with fellow family historians who are embracing the future of genealogy research!
New episodes every Tuesday. Subscribe so you never miss the latest AI tools and techniques for family history research.
I'm not going up.
SPEAKER_00Three weeks ago, I thought I had her hometown nailed down. Then I found two more records with her name on them, and each one spelled the town a different way, and not one of these three spellings matched a single line in any gazetter, old or new. If you have Eastern European Jewish ancestors anywhere in your tree, I need you to hear this before you spend one more weekend hunting for a town that might not exist under the name you were searching for. Today I'm walking you through the tool combination that gets you further than a correct spelling ever could, plus a nearly 40-year-old piece of math built specifically for this exact problem decades before anybody said the words artificial intelligence out loud. And I'm going to tell you exactly where it stops working. Let's dive in. Welcome to Ancestors and Algorithms, where family history meets artificial intelligence. I'm your host Brian, and today we're heading into Eastern European Jewish genealogy, one of the topics you all request more than almost anything else in my inbox. Now, if you don't have a drop of Jewish ancestry yourself, stick around anyway, because the actual problem we're solving today, a hometown that got spelled a different way every single time someone wrote it down, shows up in Irish research, in Italian research, in practically every immigrant story on this show. What's different here is the toolkit. Eastern European Jewish genealogists have been wrestling with exactly this problem for longer than almost anyone, out of sheer necessity, and along the way they built some of the sharpest tools anywhere for solving it. I want you to have them too, whatever part of the map your own family came from. So let's get started. Let me introduce you to Kaya Friedman. Kaya is a composite built from the kind of paper trail that shows up constantly in Eastern European Jewish research. Let's get into her story. Kaya was born around 1889 in the Pale of Settlement, the strip of Russian Empire where Jewish residence was legally restricted, and by the time she was eighteen, staying where she'd been born had stopped feeling like a choice. Part of that was legal. The May laws of 1882 had already pushed Jewish families out of the rural areas and into a shrinking number of towns, tightening what work they could do and where they could live decade after decade. Part of it was violence. In April 1903, a pogram in the city of Kishinev left dozens dead over three days and made headlines around the world, and it helped trigger a wave of immigration that would eventually carry more than a million and a half Jews out of the Russian Empire before the First World War even started. I want you to sit with what those May laws actually meant day to day, because it's easy to read legal restriction and picture something abstract. It meant a family that had farmed a piece of land for two generations could be told, with no warning attached to a calendar, that Jewish families weren't permitted to hold agricultural leases there anymore. It meant entire trades closing to you by decree, not by competition. It meant a shrinking map of towns where you were even allowed to sleep for the night. None of that shows up on a ship manifest. All of it is the reason there was a manifest to fill out at all. Kaya's older brother, Valvelle, had already made the crossing two years earlier, one thread in a much bigger pattern historians call chain migration. One sibling goes first, finds work and a place to live, and sends word back that it's safe to follow. Valvelle's own arrival isn't the story I'm telling today, but it's worth naming, because almost nobody in this wave of Eastern European Jewish immigration crossed alone. They crossed in relays, one name pulling the next one after it. A pattern that shows up so often in ship manifest from this era that joining a relative became its own printed column on the form. In the spring of nineteen oh seven, Kaya followed him. Here's what the record actually shows, not what family stories usually claim. Her ship's manifest, filled out by a clerk at the port of departure and carried across the ocean with her, says she was eighteen, single, a seamstress, and in possession of nine dollars. It says she was going to join a relative and it names him, her brother Velville Friedman, at an address on Ludlow Street on New York's Lower East Side. Nine dollars and a brother's address. That's what she was holding on to when the boat left the dock. I want to clear something up right here, because it matters for everything that comes next. You've probably heard that immigration officials at Ellis Island changed people's names, wrote down something easier to spell, and sent people off into America carrying a name that was never really theirs. I need to tell you, as plainly as I can, that this did not happen. Immigration inspectors at Ellis Island never wrote a single passenger's name down for the first time. They worked from the manifest the shipping line had already prepared back in Europe, and their actual job that day looked a lot more like fact checking than inventing. When a name did change, and plenty of them did, it changed later, at the hands of the immigrant themselves, or a naturalization clerk spelling things by ear years down the road. Kaya's own paperwork proves it. She arrived as Kaya Friedman. She left Ellis Island as Kaya Friedman. The records agree on that much completely. Where the records stop agreeing is her hometown, and that's the wall I actually want to spend today on. I know some of you have been exactly here, a name in an old document or scrawled in a margin somewhere that doesn't match anything you can find on a map, current or historical. You search the spelling you have, nothing comes back. You try a slightly different spelling, still nothing. Eventually you start to wonder if the place ever existed at all, or if somebody got the spelling wrong entirely. That's not a failure of your searching. It's the single most common wall in this entire branch of genealogy, and it's exactly why I built AI as your research assistant, not your researcher, into the name of the show in the first place. A search engine wants an exact match. A person or the right AI tool used the right way can work with something messier, a name that sounds roughly right, spelled by three different clerks who'd never heard it before and were guessing at the sounds coming out of a stranger's mouth in a language most of them didn't speak. Kaya filed her own petition for naturalization in 1917, unmarried and independent enough under the law of the time to do that entirely on her own. Between that 1907 manifest and the 1920 census still to come, spanning 13 years, her hometown would end up spelled three different ways across three different documents in three different hands, and not one of those spellings would return a single result if you typed it into a map search today. So here's the question I'm chasing today. Can any AI tool, or any combination of them, find a shtetl that three different government clerks couldn't even agree on how to spell? Let's find out. Before I trust a single one of these three spellings, I need to know I'm even reading them correctly. Ship manifest handwriting from 1907 was not built for readability, and a clerk writing quickly at the end of a long shift didn't write for people like me squinting at a scan more than a century later. This is exactly the kind of job Gemini has gotten remarkably good at over the past year, and I mean that as a verified current claim, not a vague compliment. As of August 2026, independent testing on 18th and 19th century English language handwriting has clocked Gemini's character error rate down around 1 to 2%. Expert transcriptionist territory, a huge jump from where these tools sat even a year earlier. Genealogists have also started testing it specifically on Russian language handwriting, since so much Eastern European record keeping from this era was done in Russian, and early results there are promising too, though nobody has benchmarked Russian language cursive with the same rigor yet. I want you to see exactly what I did, because you can run this on your own hardest to read document tonight. Prompt Gemini via Google AI Studio. Quote, here's a scan line from a 1907 ship manifest. Transcribe exactly what is written in the column for the passenger's last permanent residence, letter by letter, without correcting the spelling to match a place you think you recognize. If any letter is genuinely ambiguous, tell me which one and what the alternatives could be rather than guessing silently. Here's what came back. Jim and I read the line as K R I N I T Z with one flag. The second letter could be a R or could be a cramped R H combination, the kind of thing a tired clerk's pen does at the end of a word he'd never had to spell before. Not a clean, confident transcription. An honest one with the uncertainty built in instead of smoothed over. That flag matters more than it might sound like. If I'd asked a tool that only guesses at the nearest real place name instead of transcribing what's actually on the page, I might have gotten a confident wrong answer instead of an honest, slightly uncertain one. Reading the ink first before you go looking for what it means isn't a small step. It's the whole game. So now I had three documented spellings, each transcribed carefully rather than assumed. Krinitz from the nineteen oh seven Manifest, Krinitzov from the nineteen seventeen Naturalization Petition, Krinits from the nineteen twenty census where the enumerator also recorded her mother tongue as Yiddish. That question mattered. The nineteen twenty census specifically asked for mother tongue because so many immigrants from this exact region kept getting lumped under whatever empire happened to be ruling their town that decade. The language someone actually spoke at home told you far more about where they were really from than a shifting political border ever could. Three spellings, three documents, three different clerks, and not one of them writing down a place they'd ever heard of before that day. That's worth pausing on because it's easy to blame the family for a confusing name when the actual problem sat entirely on the other side of the desk. One more detail survives across the documents. In her own words to the naturalization examiner, Kaya described her hometown as being near Slonum, a real town in Grodno Province with its own substantial Jewish community and its own well documented records. Near Slonum, not Slonum itself. That distinction is about to matter a great deal. So I did the direct thing first. I went straight to Slonum's own vital records, indexed through the region's Jewish genealogical societies, and searched for a Friedman family matching Kaya's parents, names, and roughly the right years. Nothing. Not a partial match, not a maybe. That's a dead end. And I want to sit with it for a second instead of rushing past it, because this is exactly the moment a lot of research stalls out for good. Near slonum doesn't mean slonum. It means somewhere close enough that a person filling out paperwork decades later, in a second language, in a hurry, reached for the nearest big name on the map instead of the tiny place she actually meant. That's not Kaya being careless. That's how people describe home when the person writing it down has never heard of the real place. Here's the pivot, and here's where the actual math starts doing some heavy lifting. Instead of searching for an exact spelling or even a nearby big city, I needed a way to search for anything that sounded roughly like Kreinitz, Kreidnitsa, or Krinitz inside the specific province where Kaya said she was from. The tool already existed and it's nearly forty years older than anything we usually talk about on this show. In nineteen eighty five, a genealogist named Gary Mokatov, later joined by Randy Deitch, built a phonetic matching system. He built it specifically because the Standard American Sound Dex, the one built into the National Archives owned census indexes, kept failing on Eastern European Jewish names. Standard Soundex was tuned for English sounds. It treated a W and a V as different letters, for instance, when across this part of the world they were routinely interchangeable on the page. The Deich Mokotov Soundex system fixed that. It still lives today inside the free Jewish Gen community's database built for exactly this kind of search. You don't need to understand the underlying math to use it. You type in a spelling and tell it which province or district to search within if you know one. It hands back every community on record whose name resolves to a similar phonetic code, ranked and filterable. You can even filter by roughly how many Jewish residents that community had in the closest surviving census. That last filter turned out to matter more than I expected. Now here's where the golden rule earns its second mention of the episode, because this is a verification moment, not only a discovery one. AI is your research assistant, not your researcher. That line matters right here because I could have handed this entire puzzle to a chatbot and asked it to hand me an answer, and it would have given me one, confidently, whether or not it was actually right. Instead, I ran the Sound Dex search myself through Jewish Jin's own tool and got back a list of real candidate communities in Grodno Province. Only then did I bring AI in to help me reason through what that list actually meant. The SoundX search returned a short list of real Jewish communities within Grodno Province that phonetically clustered near Kreinitz, Kreinitza, and Krenitz. Two possibilities stood out immediately for very different reasons. The first was Kreinki, a real market town where the eighteen ninety seven Imperial Census counted three thousand five hundred and forty two Jewish residents, seventy-one percent of the town's total population, itself under five thousand. It had its own synagogue and its own community council called a kahal. Its documented economic life was built around tanning and leatherwork, the kind of trade that shows up again and again in Jewish occupational records from this district. The remaining candidates were two much smaller villages that happened to share a similar sounding name, the kind of coincidence this whole region is full of once you actually go looking. Both had barely a handful of Jewish households recorded as of late of the 1921 census, out of populations that ran overwhelmingly Catholic or Orthodox. One of them had exactly two Jewish residents on record the year the census was taken. That's not a sheddle a family tell stories about for three generations. That's a place a single household happened to live for a while. This is where I brought Claude in. Comparing three transcribed spellings against a short list of real candidates and weighing which one actually fits everything else the records say about Kaya's family is exactly the kind of structured reasoning task this tool handles well on this show. Prompt for Claude. Quote I have three spellings of a hometown for the same immigrant transcribed from three separate documents. Krynitz 1907 Ship Manifest, Krynitsa 1917 Naturalization Petition, and Krinitz 1920 Census Mother Tongue Recorded as Yiddish. She also independently described her hometown as near Slonom in Grado Province, Russian Empire. I have three candidate towns from a Deitch Mokatov SoundX search of that province. Krainki, a market town with 3,542 Jewish residents, 71% of the town, per the 1897 census, and its own Cajal and two smaller villages within a similar sound X code, but only a handful of Jewish households as of the 1921 census. Based on population plausibility and the phonetic pattern, which candidate is best supported and what would it actually take to confirm it? Here's what came back, and I want you to hear the reasoning, not only the conclusion. Crankey is the far stronger candidate on population grounds alone. A family whose story, passed down through paperwork and memory, kept landing near a place with an active Jewish community and its own institutions. That makes more demographic sense than a family from a hamlet with five or six Jewish households, a place unlikely to have generated enough independent immigration to be worth naming at all. Let alone naming it three separate times across thirteen years. The phonetic match holds for all three candidates about equally well, so the spelling alone can't break the tie. What would actually confirm it? A document naming Kaya or her parents specifically in connection with Cranky itself. Something like a birth record, a Cajal register entry, or a page from a community list, the kind of name search, not a play search, could surface. Notice what Claude didn't do. It didn't tell me Cranky was correct. It told me Cranky was the better supported guess, and it told me exactly what kind of document would turn a guess into an actual fact. That's the genealogical proof standard doing its job right here, correlating what the evidence can support against what it can't yet prove. So I went looking for that document. This is where I brought in a third tool. By this point I had three transcript documents, a sound dex generated shortlist, and background material on cranky itself scattered across a half a dozen browser tabs. I needed something built to hold all of it at once and tell me, with citations, what it actually adds up to. Gemini Notebook, formerly known as Notebook LM, does one thing especially well. It will not tell you anything that isn't traceable back to a source you actually gave it. That's a different job than the one Claude was doing for me a moment ago. Claude reasoned across the evidence and made an argument. Gemini Notebook's whole design philosophy is the opposite of that. No argument, no outside knowledge, nothing except what's sitting in your own uploaded sources with a citation attached to every sentence it hands back. As of August 2026, the free tier holds up to 50 sources in a single notebook. It has also grown a deep research mode that can go out onto the open web and pull in cited material on its own. I deliberately didn't use that feature here. I wanted this particular check grounded only in documents I'd already verified myself, nothing pulled in fresh and unvetted. So I uploaded four sources, the three transcribed records, and a page of background on Crankie's Jewish community. Then I asked one direct question. Prompt for Gemini Notebook. Summarize what the sources do establish about her hometown and cite each claim to its source document, end quote. Here's exactly what it told me. No more and no less. None of the four uploaded sources contain a direct named connection between Kaya Friedman and Cranky. What the sources do establish, with a citation to each, her stated hometown was spelled three different ways across three government records. She independently described it as near Sloanum. And Cranky is a documented Jewish community in the correct province that phonetically matches all three spellings. No source closes the gap between the best supported guess and a confirmed fact. I want to be honest about something else, because it belongs in this story and I'm not going to rush past it to get to the next prompt. Cranky was a real place, full of real families right up until the Second World War reached it. Like so many small Jewish communities across this exact stretch of the old Russian Empire, its own community and much of what it kept did not survive what came after. That's not a footnote to why so many searches like this one hit a wall. It's a large part of the actual reason. The records that might have named Kaya's family specifically, a synagogue register, a burial record, a community list from the right decade may no longer exist. Not because nobody thought to keep them, because of what happened to the place that kept them. That's not a comfortable thing to say out loud on a podcast about finding answers. But it's the truthing, and I'd rather tell you the truthing. So where does this leave Kaya, and where does that leave you if you're staring at your own three spelling wall tonight? I'm not gonna dress this up as a win, because a partial answer wearing a breakthrough's clothes helps nobody. Kaya's hometown is not confirmed. What I have is a strong, honestly qualified, best supported candidate. Cranky, a real Jewish community in Grodno Province. It's matched on population plausibility and phonetics across three independently transcribed spellings and ranked well ahead of two much smaller, much less demographically likely neighbors. That's real progress. It's also, and I want to say this plainly, not a name I'd put in a family tree with any confidence yet. Here's the soundly reasoned conclusion the evidence actually supports. A young woman named Kaya Friedman left the Pell of Settlement in nineteen oh seven, arriving with nine dollars and her brother's address. She left in the wake of a wave of violence that pushed over a million and a half people out of the Russian Empire in a single generation. Three separate documents, filled out by three separate hands over thirteen years, could not agree on how to spell the name of the place she'd left. The strongest candidate for that place is Cranky, a real town with a real Jewish community. Whether Kaya's specific family came from Cranky itself or from one of a small number of similarly named places nearby remains genuinely unknown, and the records that might settle it may not have survived the century that followed. I keep coming back to one detail on that ship manifest, the one that stuck with me since I first read it. Nine dollars. Not much of anything, even in nineteen oh seven, and a brother's name and address, memorized or carried on a scrap of paper, that was the whole plan. Not a fortune, not a guarantee. Enough to get to a door somebody she trusted would open. That's the human want sitting underneath every spelling variant, every sound ex code, every dead end. Kaya didn't leave home because a town's name was hard to spell. She left because home had stopped being safe, and because somewhere across an ocean, a door was going to open. The genealogy, the tools, the forty year old math problem hiding inside a modern AI workflow, all of it is how we go looking for the door behind the door. Sometimes we find the exact address. Sometimes, like today, we find the right neighborhood and have to be honest that it isn't quite the same thing. Your homework this week doesn't require a sound X table or an Eastern European ancestor. It requires one place name in your own tree that has never quite matched a map, no matter how you've spelled it. Pull every document you have that mentions that place. Transcribe each one carefully instead of trusting your first read, and look hard at whether they actually agree with each other. If they don't, you may not have a wrong record. You may have exactly the kind of wall Kayaz's. Real, honest, and worth one more look with the right tool instead of the same search you've already run four times. And for my Australian and UK listeners, this same wave of immigration reached you too, not only New York. Melbourne and Sydney both built substantial Jewish communities largely out of Eastern European arrivals from this exact period. London's East End did the same on even a larger scale, absorbing tens of thousands of people fleeing the same pill of settlement Kaya left. If a hometown in your own family's paperwork has never matched a map, the Deitch Mokata Sound Dex search works from anywhere in the world, since Jewish Gen's database is free and online no matter where you're logging in from. For the research that comes after you've narrowed a place, Australian researchers should check the National Archives of Australia at naa.gov.au. UK researchers will find the National Archives at nationalarchives.gov.uk invaluable for tracing a family once they reach your shores. Same approach, different archives, same forty-year-old algorithms doing the quiet work underneath either search. Thank you so much for listening to Ancestors and Algorithms. If you enjoyed this episode, please leave a review wherever you listen to podcasts. And if you know a fellow genealogist who could benefit from what we cover today, share this episode with them. That is the best way to help our community grow. Now, if you want to go even deeper, I've put together a 40-page companion guide for this episode over at ancestorsai.com. It includes the full Deitch Mokotov walkthrough with worked examples beyond cranky, advanced prompts for correlating conflicting spellings across any language, and a research checklist built specifically for Eastern European Jewish genealogy. But whether you grab that or not, everything we covered today gives you a solid foundation to start right away. For everything you need, including every episode, our private Facebook community, companion guides, and the research lab, head over to ancestorsnai.com. It's all right there waiting for you. I'm your host Brian, and I will see you next week for another journey into the past powered by the future. Until then, happy researching!