<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Pixels and Patterns]]></title><description><![CDATA[Research and reflections at the intersection of social science, digital humanities, and artificial intelligence.]]></description><link>https://www.pixelsandpatterns.org</link><image><url>https://substackcdn.com/image/fetch/$s_!YGtb!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5c1c326-7cd0-4791-ad50-7d68a67a6453_512x512.png</url><title>Pixels and Patterns</title><link>https://www.pixelsandpatterns.org</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 00:56:29 GMT</lastBuildDate><atom:link href="https://www.pixelsandpatterns.org/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Steven Denney]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[pixelsandpatterns@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[pixelsandpatterns@substack.com]]></itunes:email><itunes:name><![CDATA[Steven Denney]]></itunes:name></itunes:owner><itunes:author><![CDATA[Steven Denney]]></itunes:author><googleplay:owner><![CDATA[pixelsandpatterns@substack.com]]></googleplay:owner><googleplay:email><![CDATA[pixelsandpatterns@substack.com]]></googleplay:email><googleplay:author><![CDATA[Steven Denney]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Talk to your terminal]]></title><description><![CDATA[Dictating prompts to a CLI agent beats typing them.]]></description><link>https://www.pixelsandpatterns.org/p/talk-to-your-terminal</link><guid isPermaLink="false">https://www.pixelsandpatterns.org/p/talk-to-your-terminal</guid><dc:creator><![CDATA[Steven Denney]]></dc:creator><pubDate>Mon, 13 Jul 2026 12:07:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2b2dbf0a-f9e1-479c-a1f5-a52b699f4a2f_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8LRc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8LRc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8LRc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2183605,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/206812575?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8LRc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!8LRc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42726d8e-8453-4ba2-b979-8832bad4b46c_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most of my work with a command-line agent is writing prompts. Claude Code and Codex do best when I can provide as much information as possible. They excel at determining what I want from the latent constructs in the text I provide. Detailed prompts of this kind easily run three hundred words long (and longer), and typing three hundred words for every prompt, all day, is a kind of energy-sapping drudgery. The rational response is to shorten the prompt, which also degrades the answer and the work being done.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p>Earlier this year, I discovered just how good the speech-to-text technology is and created setups for my desktop machine which runs on Linux and my Macbook Pro. I stopped typing most of my prompts and started <em>dictating</em> them. I tap a hotkey and talk at the terminal through my mic for as long as the thought takes. Another tap ends the recording, and about a second later the cleaned-up text is sitting at my cursor. Accordingly, my prompts got longer and the answers and outputs got better, because a language model is very good at pulling intent out of even meandering speech. It needs volume and specifics, and speaking is the cheapest way to supply both. I still type anything that must be exact, such as slash commands, file paths, flag names, and model IDs, but almost everything else I say out loud.</p><p>This configuration has been especially useful for work that benefits from repeated explanation and close supervision. I use it to iterate on my own writing, draft emails, develop prose from rough ideas, and produce detailed feedback on student assignments, particularly writing assignments. In that setting, the limiting factor is often not what I have to say, but how much of it I am willing to type. Thorough feedback requires explanation: identifying the problem, showing why it matters, suggesting a revision, and sometimes distinguishing between several possible ways forward.</p><p>The advantage is not simply that speaking is faster. It also makes it easier to think through a response in full. I can talk through an argument, qualify a judgment, add examples, and revise my own assessment as I go. I then clean it up through edits, mostly typed but also some spoken. For student advising, this has been a big improvement. I can give more specific and more useful feedback. I have not subjected this to an A/B test, but my impression, and I am open about this with my students, is that the feedback is <em>much</em> <em>better</em> than what they are used to receiving. </p><p>Typing, however, is not obsolete. It remains better for anything that must be exact. The most effective workflow combines the two. As I will explain below, there is some modification of your voice as it is produced as text, especially if you want it to be clean. This post documents the improvement in the underlying technology, then the setup I run on Linux and macOS. Both configurations are public, <a href="https://github.com/scdenney/hyperwhspr">hyperwhspr</a> for Linux and <a href="https://github.com/scdenney/macwhspr">macwhspr</a> for the Mac. Speech-to-text is useful for all sorts of cases, but my reflections from here on are based primarily on my experiences from interacting with Agentic AI.</p><h2>Speech recognition is good now</h2><p>For a long time, dictation software produced enough errors that correcting the transcript took longer than typing it would have. In 2011, Microsoft researchers cut the word error rate on conversational telephone speech from 27.4 percent to 18.5 percent by swapping Gaussian mixture models for deep networks (<a href="https://www.isca-archive.org/interspeech_2011/seide11_interspeech.html">Seide, Li, and Yu 2011</a>). A system that still got nearly one word in five wrong counted as the breakthrough of its day. By 2016 the same benchmark was at 5.8 percent, alongside a measured 5.9 percent for professional human transcribers (<a href="https://arxiv.org/abs/1610.05256">Xiong et al. 2016</a>). A year later it was 5.1 percent (<a href="https://arxiv.org/abs/1708.06073">Xiong et al. 2017</a>).</p><p>OpenAI&#8217;s Whisper made high-quality speech recognition widely accessible. Trained on 680,000 hours of audio, it transcribed LibriSpeech at 2.7 percent error zero-shot and held up on accents and noise where earlier models collapsed (<a href="https://arxiv.org/abs/2212.04356">Radford et al. 2022</a>). In 2023 it went behind an API at $0.006 USD per minute (<a href="https://openai.com/index/introducing-chatgpt-and-whisper-apis/">OpenAI 2023</a>), which is 36 cents per hour of speech. The 2025 generation, gpt-4o-transcribe, beats Whisper across the FLEURS language benchmark by OpenAI's own evaluation (<a href="https://openai.com/index/introducing-our-next-generation-audio-models/">OpenAI 2025</a>). These figures come from different test sets. They are not directly comparable, but together they illustrate the broader improvement in accuracy, quality, and cost.</p><p>Once transcription accuracy became good enough, the remaining irritation was the pause after speaking. The 2026 releases address this, with streaming models that transcribe while you talk instead of after you finish. ElevenLabs claims under 150 milliseconds for Scribe v2 Realtime (<a href="https://elevenlabs.io/blog/introducing-scribe-v2-realtime">ElevenLabs 2025</a>), and OpenAI's gpt-realtime-whisper streams over a WebSocket at $0.017 per minute (<a href="https://platform.openai.com/docs/guides/speech-to-text">OpenAI docs</a>). I switched both of my machines to it this week, with notable improvements. </p><h2>The setup</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8pP7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8pP7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 424w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 848w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 1272w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8pP7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png" width="1456" height="600" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:600,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:131785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/206812575?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8pP7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 424w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 848w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 1272w, https://substackcdn.com/image/fetch/$s_!8pP7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d6cd504-4113-4852-9783-aa9141c7b7ea_1456x600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The pipeline. Configuration at scdenney/hyperwhspr (Linux) and scdenney/macwhspr (macOS).</figcaption></figure></div><p>The pipeline is the same on both machines. A hotkey starts recording. Audio streams to the transcription model while I speak. A second tap commits the buffer, and the final transcript comes back in about a second. A small LLM then repairs the transcript, and the result is pasted at my cursor in whatever window has focus, which is usually a terminal and occasionally a draft like this one.</p><h3><em>hyperwhspr</em>, for Linux</h3><p>On Linux the recorder is <a href="https://github.com/goodroot/hyprwhspr">hyprwhspr</a>, and <a href="https://github.com/scdenney/hyperwhspr">scdenney/hyperwhspr</a> is the reproducible configuration pattern I run under Omarchy/Hyprland. It holds the streaming backend config, the cleanup hook with its constrained prompt, the vocabulary file, the systemd service with its Wayland fixes, and the calibration command that uses the cleanup log to make new vocabulary rules. The README keeps a dated known-issues table, including this week's switch to the streaming model and other notes added by AI agents or by me.</p><h3><em>macwhspr</em>, for the Mac</h3><p>There is no hyprwhspr on macOS, so <a href="https://github.com/scdenney/macwhspr">scdenney/macwhspr</a> is the whole setup in one small repo. An idempotent installer copies everything into place. A small Python daemon records audio with SoX and streams it to the transcription service. Pressing and releasing the Globe key (my hotkey of choice) starts or stops recording, with Karabiner and Hammerspoon handling the keyboard shortcut and a small on-screen status indicator. launchd keeps the daemon running in the background, and if a transcript ends up in the wrong window, Ctrl-Cmd-V opens a list of the last twenty recordings so it can be recovered and then pasted where you want. The repository contains the complete setup. Aside from the platform-specific content, the transcription, cleanup, calibration, and logging pipeline is the same as on Linux.</p><p>The cleanup stage is an underrated part of this setup that I did not quite appreciate until setting it all up. The transcription model returns raw text with filler words, false starts, and little punctuation. A tightly constrained GPT-4.1 mini prompt restores punctuation, paragraph breaks, and consistent formatting, but it is explicitly told not to change the meaning or answering the content. If I dictate a question, I get back a well-formatted question, not an answer. Every raw and cleaned transcript pair is logged, and a calibration command uses those logs to update a vocabulary file with preferred spellings, names, romanizations, and other recurring corrections &#8212; this is how you can customize your use. Over time, the system gradually adapts to the way I speak.</p><p>Below is a recording loop example from my Linux machine. Hotkeys start the recording, the pill streams the live transcript while I speak, and a second pushing of the hotkeys send the cleaned-up text to the Claude Code prompt.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;6b446456-3af8-47bd-b84c-92bd27bc5f2b&quot;,&quot;duration&quot;:null}"></div><h2>Latency differences</h2><p>Until this week both machines uploaded the finished recording after I stopped talking, which meant the wait grew with the length of the dictation. The streaming backend does the work during the speech itself.</p><p>I measured both backends on identical audio, synthetic speech streamed at real-time pace over the same network. On a 6.5 second clip, batch took 1.7 seconds after stop and streaming took 0.9. On 94 seconds of audio, batch took 5.9 seconds and streaming took 0.9 again. Across the lengths I tested, the streaming wait was effectively flat at about a second. My most recent log has a 128 second dictation, a spoken complaint about a LaTeX table layout in a manuscript &#8212; a commonly written and voiced / yelled complaint, I am sure. The transcription wait alone was five seconds. Under the streaming backend, a dictation of that length now finishes in about one second. For a workflow built on dozens of dictations a day, that dead time adds up. The newer model is <em>much </em>better, in other words.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IB0x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IB0x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 424w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 848w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 1272w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IB0x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png" width="1456" height="840" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:840,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:96109,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/206812575?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IB0x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 424w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 848w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 1272w, https://substackcdn.com/image/fetch/$s_!IB0x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64446137-3b31-4b2c-850f-3dbf371c9b35_1456x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Measured 2026-07-13, one run per cell. The batch wait grows with audio length. The streaming wait stays at about a second.</figcaption></figure></div><p>Each cell in that chart is a single run, though the production logs agree with the pattern. Streaming costs $0.017 USD per minute against $0.006 for batch, so a two-minute dictation costs about three cents. The realtime model accepts no vocabulary prompt, so domain terms ride on the cleanup stage instead. A WebSocket connection has more failure modes than a file upload, and both repos keep the batch path one config line away. The last caveat is that the audio leaves the machine. If that is unacceptable, hyprwhspr supports a local ONNX backend, at the cost of about 5.7 GB of VRAM and maintaining the model yourself. I found the local option usable, but not good enough to justify those additional demands.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><h2>What a month of this costs</h2><p>The question of cost comes up whenever I describe this setup, so before writing this post I pulled thirty days of logs from both machines. Between June 13 and July 13 I dictated 662 times across the two machines, which comes to 7.7 hours of speech and 54,796 words. The transcription bill for the month was $2.77 USD, and the cleanup model added another $0.27, so the total was $3.04, or about five and a half cents per thousand words. The whole window ran on the older batch backend, and at streaming prices (my latest setup) it would come to about $8.11. For me, the difference is negligible. The deciding factor is the latency.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qyQ6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qyQ6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 424w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 848w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 1272w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qyQ6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:111159,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/206812575?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qyQ6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 424w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 848w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 1272w, https://substackcdn.com/image/fetch/$s_!qyQ6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c2c8ef-bcd5-419d-a4d5-5a582ac2ce77_1456x880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Thirty days of dictation, stacked by machine. Mac minutes are measured per recording, and Linux minutes are estimated from each day&#8217;s words at the calibrated pace. All of it ran on the batch backend.</figcaption></figure></div><h2>Make the switch</h2><p>Typing still has an imporant role, just to be clear. It is better for writing commands, file paths, quotations, citations, identifiers, and anything else that must be exact. It also remains an important part of writing itself. I am not about to dictate an entire scientific manuscript. Drafting, revising, and polishing still belong at the keyboard &#8212; at least for me, for now.</p><p>But typing should not be the only way you work with AI agents. Restricting yourself to the keyboard either limits the context you provide or makes providing that context unnecessarily slow and tiring. The better workflow is hybrid. Explain the problem out loud, let the transcription and cleanup stages produce a first draft, then edit it if needed, adding the precise commands, paths, or other details that only typing can reliably supply.</p><p>Speech recognition is accurate, inexpensive, and nearly immediate. Talking to an agent is no longer a novelty or an accessibility matter. It is a practical way to give these systems that are coming to dominate our computational and other kinds of pipelines the detailed instructions they require. If you do substantial work with agents and still type every prompt from beginning to end, you are imposing a bottleneck the technology has largely removed.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Who are the Americans?]]></title><description><![CDATA[On the 250th birthday of the United States, I reflect on thirty years of survey data to show a national identity in flux and make a bullish case for liberal nationalism.]]></description><link>https://www.pixelsandpatterns.org/p/who-are-the-americans</link><guid isPermaLink="false">https://www.pixelsandpatterns.org/p/who-are-the-americans</guid><dc:creator><![CDATA[Steven Denney]]></dc:creator><pubDate>Sat, 04 Jul 2026 15:20:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gj0L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gj0L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gj0L!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gj0L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2354793,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204947762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gj0L!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gj0L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b63f57e-f740-46c1-8a86-a43d7339bbe6_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Note: Data and replication code for the analysis provided here are <a href="https://github.com/scdenney/four-americas-replication">available on GitHub</a>.</em></p><p>The United States turns 250 this year, and the old argument over what it means to be American remains as intense as ever. What better moment than now to engage in that age-old debate of what it means to be an American? Flatten by media coverage, the debate usually gets <a href="https://www.nytimes.com/2026/06/09/opinion/america-250-national-identity.html">cast as two options</a>. Either the country is a creed, a set of ideas anyone can sign up to, or it is a <a href="https://www.nytimes.com/2025/12/17/opinion/republican-identity-divide.html">heritage</a>, as a matter of belonging that you inherit by birth and/or blood.</p><p>That two-sided framing is not merely an artifact of punditry. It dates back at least to the historian Hans Kohn, who, in 1944, distinguished between civic and ethnic nationalism. Civic nationalism defines the nation in political terms, as a community of citizens bound by shared institutions and ideals. Ethnic nationalism defines it in cultural and genealogical terms, as a people united by common ancestry, language, or heritage. At its core, Kohn's distinction concerned how people answered a fundamental question: <a href="https://www.cambridge.org/core/books/immigration-and-the-american-ethos/2F57F30CC2EC3110230C719EB023C111">who belongs to the &#8220;circle of we"</a>? To a large degree, this civic-ethnic distinction has organized the study of nationalism ever since. Scholars have, of course, challenged how clean the divide really is. As Anthony Smith argued, even civic nations draw on ethnic myths and symbols.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p>As public opinion surveys became a cheaper and more reliable way to measure public opinion, political scientists and sociologists began <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/nana.12947">asking people directly</a> who belongs to the nation and how they relate to it, rather than inferring those views from elite rhetoric. Deborah Schildkraut&#8217;s <em><a href="https://www.cambridge.org/core/books/americanism-in-the-twentyfirst-century/F448C226AD2607AB6E30DC295F437503">Americanism in the Twenty-First Century</a></em> and Elizabeth Theiss-Morse&#8217;s <em><a href="https://www.cambridge.org/core/books/who-counts-as-an-american/ABD0630256EFAA098D18BFFFB6AEBD0C">Who Counts as an American?</a></em> both show that Americans&#8217; views are more varied than a simple creed-or-heritage divide suggests.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>When Americans are asked directly, they do not divide neatly into two camps. Using a method that identifies recurring patterns across attitudes toward national belonging and national pride, <a href="https://wp.nyu.edu/bonikowski/wp-content/uploads/sites/18814/2020/08/bonikowski_and_dimaggio_-_varieties_of_american_popular_nationalism.pdf">Bart Bonikowski and Paul DiMaggio</a> distinguish four dispositions. Their typology, however, was developed from a 2004 survey and only later compared with data from 1996 and 2012. The question is whether those same four types still structure American national identity today.</p><p>I replicated their approach using the <a href="https://www.gesis.org/en/issp/data-and-documentation/national-identity">International Social Survey Programme&#8217;s (ISSP) National Identity module</a>, which asked the same battery of questions in 1995, 2003, 2013, and 2023. I find, again, the same four identiy types. Americans do not sort into two competing visions of the nation, but four. What has changed over the past three decades is the relative size of those groups. One has remained remarkably stable, while the others have risen and fallen.</p><h2>Two questions, four Americas</h2><p>The types follow from two basic categories of questions. First, <em>who counts as truly American</em>? Is it about being born here and being Christian, or about citizenship and respecting the country's institutions? Second, <em>how warm you feel</em> about America. How proud are you of the country? Crossing these two dimensions gives you four identity types.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OAU3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OAU3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 424w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 848w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OAU3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png" width="1456" height="1188" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1188,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:109885,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204947762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OAU3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 424w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 848w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 1272w, https://substackcdn.com/image/fetch/$s_!OAU3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63261e7d-1db2-41f4-a31f-8a9702b69133_2280x1860.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>A stylized graph of the four identity types and the questions that define them. </em></figcaption></figure></div><p>The horizontal axis captures who counts as truly American, from inclusive (citizenship and shared institutions) to exclusive (birth and religion). The vertical axis captures national pride. The types and their attitudes are then as follows:</p><ul><li><p><strong>Ardent:</strong> &#8220;hot&#8221; nationalists, scoring high on everything. They want high boundaries for belonging and are proud across the board.</p></li><li><p><strong>Restrictive:</strong> people who also draw distinctive boundaries of belonging and are sure the country is superior to others, but their pride in its actual achievements runs cooler than ardent nationalists.</p></li><li><p><strong>Creedal:</strong> those who are proud and inclusive. They anchor belonging in citizenship and institutions rather than birth or ethnicity/religion, and they are less likely than ardent or restrictive nationalists to say the country is better than others.</p></li><li><p><strong>Disengaged</strong>: those &#8220;cool&#8221; on all of it, rejecting or downplaying any sort of national identification.</p></li></ul><p>You can see each type's profile in the survey responses. The clearest divide is not national pride, which most engaged Americans share, but the boundaries of belonging. The ardent and restrictive types are much more likely to say that being Christian is part of being truly American and that the United States is superior to other countries. Creedal Americans reject both of these items as conditions for national membership. They are low on the ascriptive criteria (born here, be Christian) and on superiority, but high on the civic criteria (citizenship, respecting institutions) and moderately proud, the signature of an inclusive-but-attached patriotism. The data presented so far already complicates the creedal-versus-heritage framing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YpgI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YpgI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 424w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 848w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 1272w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YpgI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:132956,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204947762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YpgI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 424w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 848w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 1272w, https://substackcdn.com/image/fetch/$s_!YpgI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4790bc91-eeed-44ba-b667-95b2929587ed_3360x1680.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>What each identity type endorses from the pooled data (1995-2023). Each bar is the share of that type who say a criterion is important (very or somewhat) to being truly American, or that they are proud (very or somewhat) of it, colored by domain types. </em></figcaption></figure></div><h2>How the types are measured</h2><p>Since this Substack is meant to be somewhat technical, a word on how the types are constructed. The data come from the ISSP National Identity module, which fielded comparable nationally representative U.S. surveys in 1995, 2003, 2013, and 2023. I use the sixteen items included in all four waves.</p><p>As the figure above shows, the questions fall into three broad domains: who counts as truly American, how proud respondents are of different aspects of the country, and whether they express a sense of national superiority or unconditional loyalty. The first domain distinguishes civic criteria, such as citizenship and respect for institutions, from ascriptive ones, such as birthplace and Christianity. The first eleven items use four-point response scales; the remaining five use five-point agree-disagree scales.</p><p>The identity types are derived from <a href="https://journals.sagepub.com/doi/10.1177/0095798420930932?__cf_chl_f_tk=fg.3QIdpUlLf2K8udyfTNxXqEPbx6tZLFhilvbP4Jww-1783156218-1.0.1.1-867h.T_3nsyLteNfGqGAzUdh4tCRGP6qyfFfMFlYH1k">latent class analysis.</a> The method treats the population as a mixture of unobserved groups, each defined by a distinct pattern of responses. It estimates both the size of each group and the probability that each respondent belongs to it. Unlike <a href="https://www.qualtrics.com/articles/strategy-research/factor-analysis/">factor analysis</a>, a commonly <a href="https://journals.sagepub.com/doi/10.1177/000312240907400404">used alternative in the study of national identity</a>, which places respondents on one or more continuous dimensions, latent class analysis identifies qualitatively distinct types of people whose responses <em>cluster</em> together. That makes it well suited to studying national identity when the question is not how much nationalism people express, but how they combine different ideas about belonging and national pride. Bonikowski and DiMaggio adopted this approach, which is why "creedal" here refers to a type of respondent rather than a high score on, say, a creedal scale. </p><p>I estimate latent class models with one to six classes and selected the four-class solution based on model fit and interpretability. The classes are well separated, and the same four-type structure appears in all four survey waves.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><h2>What changed</h2><p>The next question is how these four conceptions of the nation have changed over time. Below is the proportions belonging to one of the four over the last thirty years. What do we see?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qoa4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qoa4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 424w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 848w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 1272w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qoa4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png" width="1456" height="881" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:881,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:161232,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204947762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qoa4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 424w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 848w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 1272w, https://substackcdn.com/image/fetch/$s_!Qoa4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8d4c94e-16a7-4f23-9b52-8f98868bafeb_2280x1380.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The share of Americans in each type, 1995 to 2023. </em></figcaption></figure></div><p>Two things stand out. First, the two hot types have receded. The September 11 terrorist attacks produced a sharp but temporary surge in more exclusionary forms of nationalism. The ardent type reached its peak in 2003 before steadily declining, while the restrictive type has fallen below its pre-9/11 level. Second, the creedal type experienced the inverse of that surge. Its share fell sharply in 2003 but returned to its long-run level of roughly one-third of Americans by 2013, where it has remained since. Across three of the four survey waves, the inclusive conception of the nation is the modal form of American nationalism.</p><p>Perhaps most striking is <em>the growth of the disengaged</em>. In 1995, they accounted for about one in eight Americans. By 2023, they had become the largest of the four groups, comprising more than one-third of the population. Because these are repeated cross-sectional surveys rather than panel data, we cannot conclude that individuals moved from the ardent or restrictive types into disengagement. But the aggregate pattern is consistent with exactly that possibility.</p><p>This finding should hold our attention. In 2023, <em>the most common way of relating to the nation is detachment</em>. It is hard not to make this connection to a broader <a href="https://www.journalofdemocracy.org/articles/bowling-alone-americas-declining-social-capital/">retreat from shared life and experiences</a>. <a href="https://www.pewresearch.org/2025/05/08/americans-trust-in-one-another/">Interpersonal trust</a> has fallen, while community and religious membership has <a href="https://news.gallup.com/poll/341963/church-membership-falls-below-majority-first-time.aspx">thinned out</a>. <a href="https://www.hhs.gov/sites/default/files/surgeon-general-social-connection-advisory.pdf">Isolation</a> has risen to the point that the Surgeon General has called it an epidemic. The problem is less that Americans hate one another's idea of the country than that a growing number are opting out of having one at all.</p><p>Politically, the next obvious question is whether this shift has unfolded similarly between left/right partisans. Looking separately at Democrats and Republicans show both encouraging similarities and notable differences.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x8Mu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x8Mu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 424w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 848w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x8Mu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png" width="1456" height="713" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:713,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:197646,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204947762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!x8Mu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 424w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 848w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 1272w, https://substackcdn.com/image/fetch/$s_!x8Mu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F787e1108-1771-4625-b954-c72ac471dbaa_2940x1440.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The four types over time within each party. Democrats (L), Republicans (R).</em></figcaption></figure></div><p>The most striking finding is <em>the stability of the creedal type across the partisan divide</em>. Throughout the past three decades, roughly one-third of both Democrats and Republicans have held an inclusive but proud conception of the nation.</p><p>Where the parties differ is in the trajectory of the other three types. Among Democrats, the hot and boundary-drawing types declined, with most of that shift moving toward disengagement, which now accounts for <em>more than four in ten</em> Democrats.</p><p>Among Republicans, the restrictive type remained relatively stable, although it now appears to be declining as well. Instead, the defining change over the past three decades was the sharp post-9/11 surge in ardent nationalism, followed by its steady decline. Disengaged Republicans also became more common, although they remain much less prevalent than among Democrats.</p><h2>The case for liberal nationalism</h2><p>The creedal conception of the nation has proven more durable than the others. For those who see liberal nationalism as the most desirable basis for democratic politics, that is encouraging news. Creedal Americans define national belonging in terms of citizenship and shared institutions rather than ancestry, while expressing pride in their country without believing it is superior to others. This is what political theorists call <a href="https://press.princeton.edu/books/paperback/9780691001746/liberal-nationalism">liberal nationalism</a>. It is, among self-identifying nationalists, a normatively desirable form of thinking about community in the nation-state era.</p><p>The central claim of liberal nationalism is straightforward. A shared civic identity can sustain liberal institutions by encouraging strangers to see one another as members of the same political community. Rather than defining the nation by ethnicity or religion, it defines membership through citizenship and shared democratic institutions.</p><p>The distinction is not merely theoretical. It also lies at the heart of one of the country&#8217;s most consequential constitutional debates: birthright citizenship. Creedal Americans place relatively little weight on being born in the United States as a condition of being &#8220;truly American,&#8221; but they overwhelmingly regard citizenship as essential. The logic of the creedal conception therefore points toward <em>jus soli</em> &#8212; citizenship, rather than ancestry, is the basis of national belonging. The Supreme Court&#8217;s recent decision In <em><a href="https://www.supremecourt.gov/opinions/25pdf/25-365_4hdj.pdf">Trump v. Barbara</a></em><a href="https://www.supremecourt.gov/opinions/25pdf/25-365_4hdj.pdf"> (2026</a>), reaffirming birthright citizenship under the Fourteenth Amendment, reflects this civic understanding of national membership.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>  </p><p>In recent years, a growing number of scholars have argued that liberals <a href="https://www.journalofdemocracy.org/articles/the-power-of-liberal-nationalism/">should reclaim</a> this inclusive form of nationalism rather than cede the idea of the nation to more exclusionary visions. Writers making that case for America, from <a href="https://www.noahpinion.blog/p/america-needs-liberal-nationalism">Noah Smith</a> to <a href="https://www.popularbydesign.org/p/why-us-nationalism-is-essential-after">Alexander Kustov</a> and others, argue that the strongest foundation for national identity is one rooted in citizenship rather than ancestry, open to newcomers, and confident enough in its ideals to welcome them as fellow citizens. Attempts to bypass national identity altogether, whether through cosmopolitanism or other forms of supra- or non-national identification, they argue, simply leave the language of nationhood to the demagogues.</p><p>My findings suggest that this civic conception of the nation is more than a normative ideal. While ardent, restrictive, and disengaged forms of national identity have expanded and contracted over the past three decades, the creedal type has remained remarkably stable, accounting for roughly one-third of Americans in every survey. Nor is that durability merely an American curiosity. In my comparative study of twenty-nine countries, the United States emerges as the clearest example of civic nationalism: the country where commitment to democratic rights most strongly constrains hostility toward outsiders.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><h2>The road ahead</h2><p>So the mood on America's 250th should be one of cautious optimism. The meaning of the American nation remains unsettled. The &#8220;<a href="https://x.com/McConaughey/status/2073013891937034425">bets still on the table</a>&#8221;. The civic conception of the country has endured, exclusionary nationalism has not become dominant, but a growing share of Americans remains disengaged rather than committed to either vision. </p><p>There is, to be sure, a real loss of attachment. But that also leaves space <a href="https://hiddentribes.us/">for persuasion</a>. As <a href="https://www.youtube.com/watch?v=5k4oajXMsVQ&amp;themeRefresh=1">Ross Douthat argues</a>, America&#8217;s enduring strength is its capacity to continually renew itself. This is a contested nation, but <a href="https://www.pewresearch.org/short-reads/2026/03/25/how-americans-value-racial-diversity-ahead-of-the-countrys-250th-anniversary/">diversity is somethig most Americans think makes the country stronger</a>, anyway.</p><p>The central challenge, then, is not simply to contain exclusionary nationalism, which is always going to exist in some form or another. It is to build a more compelling alternative. The evidence suggests that the most durable conception of the American nation is also the most <em>small-l</em> liberal. It is the one that grounds belonging in citizenship, shared institutions, and common political commitments rather than ascriptive characteristics. For advocates of liberal nationalism, the task is therefore not only to oppose exclusion, but to persuade the growing number of disengaged Americans, especially on the left, that the American experiment is still a project worth identifying with and investing in.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Hans Kohn, <em>The Idea of Nationalism: A Study in Its Origins and Background</em> (1944), is the source of the civic-versus-ethnic distinction. Anthony Smith spent much of his career complicating it, arguing that civic nations carry their own ethnic and symbolic content. Rogers Brubaker&#8217;s <em>Citizenship and Nationhood in France and Germany</em> (1992) shows the distinction operating as law, not just rhetoric. Even Samuel Huntington, in <em>Who Are We?</em> (2004), pushed against a purely creedal reading of American identity, though from the other direction. He argued the creed itself grew out of, and depends on, an Anglo-Protestant culture rather than standing free of one. Rogers Smith&#8217;s <em>Civic Ideals</em> (1997) is the fullest brief for a third view, that American political culture has always uncomfortably and contentiously bound liberal, republican, and ascriptive traditions together rather than resting on the creed alone.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Both books draw on original national surveys, not just theory. Jack Citrin&#8217;s decades of work on American identity, much of it with co-authors including Donald Green and David Sears, can be read as the earlier generation of this empirical tradition.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>The four-type structure is imposed as a modeling constraint. I estimate a multi-group latent class model that holds the class-specific response profiles constant across waves while allowing only class proportions to vary. I compared this specification with an alternative that allowed the class-specific response profiles to vary across waves. The constrained model was preferred because it had a lower Bayesian Information Criterion (BIC), a standard measure that balances model fit against unnecessary complexity. So, the class definitions are estimated jointly from all four waves. Holding the measurement model fixed ensures that changes in class proportions reflect <em>changes in the prevalence of the same latent types over time</em>. Estimating the model separately for each wave would redefine the classes in each year, making longitudinal comparisons of the identity groups confusing.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Citizenship is one of the few criteria endorsed across all four types: 93 percent of creedal respondents and essentially all restrictive respondents say it is important to being truly American. Birthplace is far more divisive. Between 78 and 91 percent of ardent and restrictive respondents say one must be born in the United States, compared with 45 percent of the creedal type. <a href="https://www.pewresearch.org/short-reads/2021/05/25/in-both-parties-fewer-now-say-being-christian-or-being-born-in-u-s-is-important-to-being-truly-american/">Pew likewise finds</a> declining support for birthplace as a criterion of American identity, especially among Democrats. Notably, roughly <a href="https://apnorc.org/projects/only-a-quarter-believe-that-the-u-s-is-a-great-place-for-immigrants/">two-thirds of Americans support birthright citizenship</a>, although support drops when surveys ask specifically about children born to parents who are in the country unlawfully.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>I show across 29 countries that democratic commitment restrains the exclusionary face of nationalism mainly where nationhood is imagined in civic rather than ethnic terms. The US is the civic nation exemplar, with the strongest such restraint of any country in the sample. See: Steven Denney, Democracy and Nationalism, Reconsidered (June 09, 2026). Available at SSRN: <a href="https://ssrn.com/abstract=6904542">https://ssrn.com/abstract=6904542</a>.</p></div></div>]]></content:encoded></item><item><title><![CDATA[From hallucination to verification]]></title><description><![CDATA[AI has made fabricated citations commonplace. Used correctly, AI can make them rare. I show you how, using "skills".]]></description><link>https://www.pixelsandpatterns.org/p/from-hallucination-to-verification</link><guid isPermaLink="false">https://www.pixelsandpatterns.org/p/from-hallucination-to-verification</guid><dc:creator><![CDATA[Steven Denney]]></dc:creator><pubDate>Tue, 30 Jun 2026 07:23:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rk-W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rk-W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rk-W!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rk-W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2239327,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.pixelsandpatterns.org/i/204137437?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rk-W!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rk-W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6f06942-85a9-44c9-8dae-5fe2981c633e_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Earlier this month, <a href="https://x.com/ChrisCarothers/status/2065792749282943030">a widely shared thread on X</a> pointed out that a prominent political scientist had a half-dozen fabricated citations in a recent article on South Korean politics.<sup><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></sup></p><p>The reaction was about what you would expect: disappointment, mostly, and, in some quarters, shock. To be sure, no instructor would accept this in a student paper, and fabricated references in a published article must be addressed, whether through a correction or a retraction. Given the number of references that appear to have been fabricated, I suspect that most people following the case expect the latter.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p>Still, the speed and intensity of the response are worth pausing over. A case like this deserves more care than it has received so far.</p><p>First, we do not actually know how those citations got there. It is unlikely but entirely possible that the fault lies with the journal rather than the author. Many of us have concluded that this particular outlet does not meet the standards it advertises, and here it was turning an article around quickly to address fast-moving events in Korea. Under that kind of pressure, it is not far-fetched that an editor ran the manuscript through an AI tool to tidy or reformat the references and that the tool quietly rewrote them into works that do not exist. I do not know that this happened, and at this point I would have expected this information to have been brought to light, but neither does anyone else.</p><p>Calling out errors of this kind is entirely fair. But a public accusation of malpractice directed at an individual had better be well founded, because once it is made, it cannot easily be undone. Reputational damage travels quickly and lands hard; any later correction, if one comes at all, rarely receives comparable attention.</p><p>This moment also calls for a dose of humility. The ordinary, almost boring parts of academic work, such as locating a source and formatting a reference, have been disrupted. The tools responsible are exceptionally good at making users feel that they have completed the work correctly when, in fact, they have not. Even otherwise careful people can be duped by their personal robot.</p><p>Much of the anger on display, I suspect, reflects a broader anxiety about what AI is doing to the academy. Before making an example of someone, particularly someone with a record of excellent work, we should be far more certain about what actually happened. We should also place the case within the wider structural conditions that made it possible.</p><p>To be clear, I do not know how this happened. The most plausible in my mind is not calculated fraud but misplaced confidence in a technology that routinely produces convincing falsehoods. That does not excuse the error. It does, however, change what we should learn from it. The structural failure may be less an individual lapse than a culture that has adopted powerful tools without establishing equally powerful norms of verification. The good news is that this is a problem we can fix &#8212; and I have some solutions to try out. But first, we need to understand how we got here, and why the machine is so good at deceiving us.</p><h2>Old habits die hard</h2><p>Type a scholar&#8217;s name and a couple of keywords into Google Scholar, and you get a list of citations. Use a few of them without reading past the abstract because you already know roughly what that person argues, and you have done one of the more mundane tasks of writing. AI is just exposing an inconvenient truth. We do not (closely) read every source we cite. Some citations are indeed used as evidence, and some are presented as a kind of performance. They signal that we know the literature or that we have cited whoever a reviewer expected or explicitly asked us to. A chatbot like ChatGPT or Claude, in these early days, feels like the same exercise. You ask for the literature on a subject or particular claim, and it gives you names you recognize along with titles that look right. The problem is that some of those titles do not exist.</p><p>None of this is new. Take one prominent example. A decade ago, before anyone could blame a model, <a href="https://retractionwatch.com/2019/09/17/columbia-historian-stepping-down-after-plagiarism-finding/">a prize-winning history of North Korea</a> was found to have cited nonexistent or irrelevant sources in at least sixty-one places. The author returned his award and left his chair.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> The much less catastrophic version of the problem is ubiquitous. It includes use of the wrong DOI, a page range off by two, or a citation to a real paper for a claim it does not actually make. Hallucinations in all but name. What AI adds is ease of use and an appearance of accuracy, which contributes to the major structural issue we are all dealing with now.</p><h2>Why the machine makes things up</h2><p>The difference between Google Scholar and an AI chatbot is that the search engine (usually) returns sources that exist. A large language model returns things that are probable. It is neither a database nor a search engine. It generates the next most likely token given everything before it, so when you ask for a citation, it produces a string outcome that <em>looks</em> like one. The string has an author, a year, a title, a journal, and a volume because that shape is overwhelmingly probable wherever citations appear. Nothing in that process checks the reference against a master list. The model is indifferent to whether the work exists, a machine version of <a href="https://link.springer.com/article/10.1007/s10676-024-09775-5">Harry Frankfurt&#8217;s &#8220;bullshit&#8221;</a> &#8212; unconcerned with the truth of what it produces. A recent <a href="https://arxiv.org/abs/2509.04664">paper from OpenAI</a> argues that the problem is more deeply embedded in these systems than we like to admit. Our training and evaluation procedures reward a confident guess over an admission of uncertainty, so models learn to answer confidently rather than abstain or admit ignorance.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>How often does this happen? In <a href="https://doi.org/10.1038/s41598-023-41032-5">one controlled study</a>, GPT-3.5 fabricated 55 percent of the citations it generated for literature reviews, and a newer model (GPT-4), while considerably better, still fabricated 18 percent.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> <a href="https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(26)00603-3/fulltext">An audit of two and a half million biomedical papers</a> found that the rate of fabricated references had risen more than twelvefold since 2023, to about 57 per 10,000 papers by early 2026. The steepest increase accompanied the arrival of new AI writing tools in mid-2024. Most of those references had already passed peer review, which is a lot to ask of reviewers who would have to chase down every source.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> The problem reaches well beyond the academy. In 2023, a federal judge <a href="https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2022cv01461/575368/54/">fined two New York lawyers five thousand dollars</a> after they filed a brief built on six judicial opinions ChatGPT had invented. The lawyer had asked the chatbot to confirm that one of them was real.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kJ_j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kJ_j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 424w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 848w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kJ_j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png" width="1456" height="835" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:835,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Quarterly rate of fabricated references per 10,000 biomedical papers, rising from about four in 2023 to nearly 57 by early 2026.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Quarterly rate of fabricated references per 10,000 biomedical papers, rising from about four in 2023 to nearly 57 by early 2026." title="Quarterly rate of fabricated references per 10,000 biomedical papers, rising from about four in 2023 to nearly 57 by early 2026." srcset="https://substackcdn.com/image/fetch/$s_!kJ_j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 424w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 848w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!kJ_j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F672c5bc9-264f-404a-b086-b7a487542b87_1803x1034.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>While we do not know how the fabricated citations in the case presented in the introduction of this post were produced, the likeliest account is the simplest one. A single prompt to a general-purpose chatbot, with no tools or documents to check against. That is the most fragile way to use this new technology, and some degree of hallucination in the response(s) is close to guaranteed. However, systems designed to retrieve and read real sources before answering do far better, even if imperfectly so. When <a href="https://doi.org/10.1111/jels.12413">Stanford tested</a> the retrieval-based legal tools that vendors marketed as &#8220;hallucination-free,&#8221; they still returned false information in something like one answer in six. But this is at least a sign of things moving in the right direction.</p><p>At the very least, we need verification systems that can find hallucinations and other errors.</p><h2>Responsible all the same</h2><p>We are accountable for what we publish. There is no better safeguard than reading our sources carefully, checking every citation, and refusing to include anything we cannot verify. Fabricated sources are unacceptable whether they come from the poor use of an LLM, a tired research assistant, or our own carelessness in formatting.</p><p>For students, that standard has to be explicit, and at my own institution it is. While I am generally unimpressed with how universities &#8212; my own included &#8212; cast AI as primarily a tool for cheating, <a href="https://www.organisatiegids.universiteitleiden.nl/binaries/content/assets/geesteswetenschappen/oer/2024-2025/oer-2024-2025-ma-algemeen-fgw-eng.pdf">Leiden University&#8217;s Faculty of Humanities</a> correctly defines fraud as any act or omission that prevents a correct assessment of a student&#8217;s knowledge. Its examples include the unauthorized use of AI software and the invention of research data.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> Its <a href="https://www.organisatiegids.universiteitleiden.nl/en/regulations/humanities/guidelines-for-the-use-of-genai-in-assessment">guidance on generative AI</a> is more direct: &#8220;GenAI generates language, not information &#8230; If asked to produce references, many LLMs will simply invent non-existent citations.&#8221; A fabricated citation misrepresents the evidentiary basis of the work, which is exactly what assessment is meant to evaluate. Under a strict reading, it is indeed fraud, and the <a href="https://www.organisatiegids.universiteitleiden.nl/en/regulations/general/plagiarism">heaviest penalty on the books</a>, exclusion from all examinations for a year, should be deterrent enough.</p><p>A field-wide response is going to be<em> </em>harder than a local one, and the distance between what these policies seek to resolve and what actually happens seems quite wide. The instinct here seems to be to police the problem out of existence. Rooted in a commitment to supporting good science, it is understandable. In mid-May, <a href="https://techcrunch.com/2026/05/16/research-repository-arxiv-will-ban-authors-for-a-year-if-they-let-ai-do-all-the-work/">arXiv announced</a> a one-year ban for authors caught submitting work with hallucinated references, with any later submission required to clear peer review elsewhere first.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> The motivation is right, but the mechanism is not. I think the verdict that the policy is <a href="https://www.insidehighered.com/news/faculty/books-publishing/2026/05/22/ban-authors-who-submit-ai-content-welcome-unenforceable">&#8220;welcome but unenforceable&#8221;</a> is the correct one.</p><p>The early evidence agrees. When one peer-review service <a href="https://reviewer3.com/live/arxiv">ran a reference checker over two hundred papers</a> posted to arXiv after the ban, more than one in four still carried a hallucinated citation. Despite threats of punishment, obviously people are using AI for research. The systems are too good, and too useful, for prohibition to hold. People will use them to write and to think, whatever the rules say. In all likelihood, messages meant to deter, to the extent that they are heard at all, are probably just driving an <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5464215&amp;__cf_chl_f_tk=HjfjFtdjV0QXvMc325Mi5U2F_jPtNMXsIuyLOKbq.o0-1782765665-1.0.1.1-w81elC9vRsNfufDg1PueKB.3dTuqrDPTOBMYHIKyyyg">underreporting of AI use because of social desirability</a>.</p><p>We need a different approach here. Teach people to use these tools well, in the open, and turn the same technology around on the problem it has exacerbated. Build verification into the places that matter, such as a journal&#8217;s submission desk, and help writers &#8212; senior authors and students alike &#8212; such that auditing a reference list, or checking whether a source actually supports the claim attached to it, becomes routine. Done properly, these tasks could be performed more systematically and reliably than they were before ChatGPT. This shifts the emphasis from telling people what not to do toward giving them something better to do.</p><h2>Skills / chatbots</h2><p>The first thing we need to do is stop treating chatbots as the default interface for serious research and writing. Instead, build (or use) skills for use with <a href="https://cloud.google.com/discover/what-is-agentic-ai">agentic AI</a>.</p><p><a href="https://support.claude.com/en/articles/12512176-what-are-skills">A skill,</a> in the agentic AI sense, is a structured set of instructions that constrains how an AI system does a task. It fixes the steps the model takes and the tools it may use. It also specifies how the model checks its own output. A one-shot prompt invites the model to be creative. A skill takes that freedom away. It can direct a model to dispatch <a href="https://cloud.google.com/discover/what-are-ai-agents">agents</a> and to look up sources in real indexes and distinguish what it actually confirmed from what it merely guessed. The fabrication problem is fundamentally a model answering from probability instead of from evidence, and a well-built skill helps close off that option.</p><p>I keep a small library of these for my own research, the <a href="https://github.com/scdenney/open-science-skills">Open Science Skills</a>. Three of them are built for the problems we are entertaining here.</p><p>The first is a <a href="https://github.com/scdenney/open-science-skills/blob/main/plugin/skills/citation-check/SKILL.md">reference checker</a> (<code>citation-check</code>). This most directly addresses the problem under consideration here. It inventories every in-text citation and every entry in the bibliography of a paper, then runs the checks on them. It confirms that each cited work exists, using programmatic indexes like Crossref and OpenAlex rather than the model&#8217;s memory. It confirms that a DOI resolves to the cited paper rather than merely to <em>some</em> paper. It cross-checks a suspicious title against the named author&#8217;s actual body of work. And it will not call anything fabricated until it has failed to find it by title and by author. What comes back sorts the serious problems from the cosmetic ones and flags the uncertain cases for a human. It is <em>not</em> perfect, but it works.</p><p>The second addresses a different, arguably much deeper problem. It is one that is much less likely to go viral on X. A citation can be real and perfectly formatted and still not support the claim it is associated with. The <a href="https://github.com/scdenney/open-science-skills/blob/main/plugin/skills/fact-check/SKILL.md">source-claim checker</a> (<code>fact-check</code>) reads the cited source and asks whether it supports the specific sentence. It flags overclaiming and direction errors. Notably, it needs the sources to read. As I&#8217;ve written the skill, you cannot just direct it to the internet or ask it to &#8220;get creative&#8221;. This takes us to the third skill.</p><p>The <a href="https://github.com/scdenney/open-science-skills/blob/main/plugin/skills/research-repo/SKILL.md">knowledge-base scaffold</a> (<code>research-repo</code>) builds a project around its source library. The original formats (e.g., .pdf) are kept in one folder while another holds the machine-readable conversions the checker reads (markdown, preferably). Each is keyed to a bibliography entry for a document you have actually filed so there is a direct connection between source and its bibliographic metadata. The claim checker refuses to run against an empty or half-built library. To confirm you have represented a source correctly, you have to actually have the source &#8212; or, where the source is not available, a file that faithfully summarizes it.</p><h2>Try it yourself</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://scdenney.github.io/ai-for-research/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0y_h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 424w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 848w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 1272w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0y_h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png" width="1456" height="479" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:479,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;AI for Research &#8212; demos, lecture slides, and how-to guides for working with AI agents and skills.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://scdenney.github.io/ai-for-research/&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="AI for Research &#8212; demos, lecture slides, and how-to guides for working with AI agents and skills." title="AI for Research &#8212; demos, lecture slides, and how-to guides for working with AI agents and skills." srcset="https://substackcdn.com/image/fetch/$s_!0y_h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 424w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 848w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 1272w, https://substackcdn.com/image/fetch/$s_!0y_h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54aa7047-f32a-41c8-a101-75c1185721b7_1800x592.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The companion hub, where I will post demos and presentation slides.</figcaption></figure></div><p>None of this is as elaborate as it might sound, and the best way to figure this stuff out is to just do it yourself. I have put up a small, self-contained <a href="https://scdenney.github.io/ai-for-research/reference-check">walk-through</a> on a synthetic manuscript seeded with deliberate errors, so you can watch the checks fire away. It is the first demo in a <a href="https://scdenney.github.io/ai-for-research">teaching hub</a> (working prototype) I am building for students and colleagues at Leiden University, with the setup guide and slides alongside it.</p><p>We are in an odd period. The technology is too useful to ignore, yet unreliable enough to embarrass or even discredit us. At the same time, we have not used it long enough to develop sound habits around it. That is a dangerous combination.</p><p>The answer is not a moratorium or a rejection of the technology. Nor should it be a hunt for the next person to make an example of on social media. We need humility about how easily these tools can fool us, some grace for those who have misused them, and stricter standards for how we use them. These positions are not contradictory. We can hold all three at once.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p><em>Correction: It has since been clarified that the article published in Korea Observer was an "annual review" reflection and was therefore not subject to peer review, which explains its quick publication timeline.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>The article appeared in <em>Korea Observer</em> 57, no. 1 (Spring 2026): 1&#8211;21, https://doi.org/10.29152/KOIKS.2026.57.1.1, a 21-page review of South Korean politics after the December 2024 martial law crisis. It was submitted on February 10 and accepted on February 21, and the journal advertises itself as a peer-reviewed quarterly. The widely shared thread on June 13 flagged six citations as fabricated, and two of the scholars named confirmed the attributed works do not exist. Two of the named authors, asked by news outlet <a href="https://www.nknews.org/2026/06/award-winning-korea-studies-scholar-accused-of-using-ai-fabricated-citations/">NK News</a>, confirmed they had never written the works attributed to them. The article published, notably, was not subjected to the journal&#8217;s peer review process, as it is a special &#8220;annual review&#8221; reflection.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Charles K. Armstrong, <em>Tyranny of the Weak: North Korea and the World, 1950&#8211;1992</em> (Ithaca, NY: Cornell University Press, 2013). The citation problems were first documented in detail by the historian Bal&#225;zs Szalontai. A <a href="https://retractionwatch.com/2019/09/17/columbia-historian-stepping-down-after-plagiarism-finding/">Columbia University investigation</a> found research misconduct, concluding that Armstrong had &#8220;cited nonexistent or irrelevant sources in at least 61 instances,&#8221; alongside plagiarism; Szalontai&#8217;s own count ran higher still. In 2017 he <a href="https://retractionwatch.com/2017/07/05/historian-returns-prize-high-profile-book-70-corrections/">returned</a> the American Historical Association&#8217;s John K. Fairbank Prize, writing, &#8220;Due to the numerous citation errors in my book, I have decided to return the prize out of respect for the AHA,&#8221; and he retired from Columbia in 2020. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang, &#8220;Why Language Models Hallucinate,&#8221; arXiv:2509.04664 (2025), p. 1: &#8220;language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty &#8230; language models are optimized to be good test-takers, and guessing when uncertain improves test performance.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>William H. Walters and Esther Isabelle Wilder, &#8220;Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT,&#8221; <em>Scientific Reports</em> 13 (2023): 14045, https://doi.org/10.1038/s41598-023-41032-5, p. 5. Across 636 citations, &#8220;55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated. Likewise, 43% of the real GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Maxim Topaz, Nir Roguin, Pallavi Gupta, Zhihong Zhang, and Laura-Maria Peltonen, &#8220;Fabricated Citations: An Audit across 2&#183;5 Million Biomedical Papers,&#8221; <em>The Lancet</em> 407, no. 10541 (2026): 1779&#8211;1781, https://doi.org/10.1016/S0140-6736(26)00603-3. The audit scanned the PubMed Central open-access subset, about 2.5 million papers and 125.6 million references, of which 97.1 million carried an identifier and could be verified. It found the rate of fabricated references rising more than twelvefold, from roughly four per 10,000 papers in 2023 to 51.3 per 10,000 in the fourth quarter of 2025 and 56.9 in early 2026, with the sharp inflection in mid-2024. The share of papers with at least one fabricated reference rose from one in 2,828 in 2023 to one in 277 in the first seven weeks of 2026. The authors cite earlier work estimating that 30 to 69 percent of LLM-generated references in biomedicine are fabricated.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Leiden University, Faculty of Humanities, Course and Examination Regulations, Art. 7: &#8220;Fraud is defined as any activity or omission carried out by a student aimed at completely or partially hindering a correct assessment of the student&#8217;s knowledge, understanding and skills, including at least the following: &#8230; c. unauthorized use of AI software; d. inventing research data &#8230;&#8221; The enumerated list does not use the words &#8220;fabricated references&#8221; verbatim, so mapping a hallucinated citation to a specific clause is an interpretation, but coverage under the general definition is clear.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Reported across Nature, TechCrunch, Times Higher Education, Inside Higher Ed, and others in mid-May 2026, and announced by an arXiv moderator and computer-science section chair rather than, at first, as codified policy text on arXiv&#8217;s official pages. The stated trigger is &#8220;incontrovertible evidence&#8221; of unchecked LLM output, with hallucinated references the main example. This is distinct from arXiv&#8217;s earlier (October 2025) requirement that review and survey articles in its CS category clear peer review before posting, which is a submission rule, not an author ban.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Vision-language models are weirdly good at Old-Korean]]></title><description><![CDATA[I stress test VLMs on Old-Korean and it works surprisingly well.]]></description><link>https://www.pixelsandpatterns.org/p/vision-language-models-are-weirdly</link><guid isPermaLink="false">https://www.pixelsandpatterns.org/p/vision-language-models-are-weirdly</guid><dc:creator><![CDATA[Aron van de Pol]]></dc:creator><pubDate>Fri, 26 Jun 2026 16:24:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e79d4085-2b31-4b01-b9a7-d157fc20d870_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p></p><p>OCR has changed substantially in the last few years. For clean, modern, printed English, the task no longer usually requires a dedicated OCR engine, a domain-specific training set, or extensive preprocessing. Current vision-language models can often transcribe such material directly from a prompt, sometimes zero-shot and sometimes with a single example. In that limited sense, printed Latin-script OCR for high-resource, modern material has become a largely infrastructural problem rather than a central research challenge.</p><p>Historical documents are still harder. Old paper, damaged scans, unusual layouts, handwriting, and older typefaces all push modern models well outside their comfort zone. But even here the progress has been striking. Platforms like <em><a href="https://www.transkribus.org/">Transkribus</a></em> routinely read European manuscripts that would have seemed hopeless a decade ago, and projects such as <em><a href="https://github.com/knaw-huc/loghi">Loghi</a></em> report similarly strong results on historical print. OCR has become much more accessible.</p><p>Korean, and by extension CJK<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>  and other non-Western scripts, is a different case, and early-twentieth-century Korean is different again. Pre-1933, Korean was written in an orthography that has since fallen out of use. It spelled words closer to how they sounded than to a fixed stem, and it kept letters the modern keyboard has dropped, among them the dot vowel arae-a (&#12685;) and a set of old consonant clusters. The spelling we read today is the Unified Orthography of 1933 (&#54620;&#44544; &#47582;&#52644;&#48277; &#53685;&#51068;&#50504;), the morphophonemic standard carried by Chu Sigy&#335;ng&#8217;s lineage in the Korean Language Society, and it became the convention only after a long argument that ran through the colonial period and resurfaced in the spelling crisis of the early 1950s.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>Modern Korean OCR remains less stable than the English and European-script cases above. On the Korean portion of the <a href="https://songjhpku.github.io/PM4Bench/">PM<sup>4</sup>Bench</a> benchmark the leading system changes with each new model release, and in the current version Gemini 3 Pro ranks ahead of the newest open model, Qwen3-VL, which does not improve on its own predecessor for this task. This should be read as a relative ranking within a weak field rather than as evidence of reliable transcription. The absolute scores remain low.</p><p>In one of my projects, I need OCR labels for about ninety-four million individual character images. Because I study the appearance of historical type rather than the content of the text, it is important that older Korean characters are transcribed as they were printed, rather than silently modernized. This post explores how well current models actually do.</p><h2>A few tests pages</h2><p>The page is from <em>Sony&#335;n</em> (&#49548;&#45380;), a children&#8217;s magazine published in 1908, and contains a chain-rhyme word game. It is printed in vertical columns read from right to left, combining (Old-)Korean <em>Hangul</em> and <em>Hanja</em> on the same page. Material like this has traditionally been difficult for OCR systems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZY-M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZY-M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 424w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 848w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 1272w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZY-M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png" width="1168" height="1206" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1206,&quot;width&quot;:1168,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2260716,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ZY-M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 424w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 848w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 1272w, https://substackcdn.com/image/fetch/$s_!ZY-M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6c32c2f-cb10-45f3-9f5f-a9d5449f2bf6_1168x1206.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The part I cared about was not whether the models could produce readable Korean. (Though even this would be helpful for many historians.) It was whether they would OCR Korean letters as printed. Before the 1933 spelling reform, Korean used forms that later disappeared or changed in ordinary writing. These include the dot-vowel arae-a (&#12685;) and several old consonant clusters. A model can read a page fluently and still normalize those forms into modern Korean. For most readers the output still looks right. For my purposes, it is wrong.</p><p>The first page was a narrow test. It happened to contain two old cluster forms, &#4397;&#4449; and &#4399;&#4465;. A modern transcription would normally write these as &#44620; and &#46832;. GPT-5.5 kept all six occurrences of those old clusters on every run. The previous GPT model had modernized all six. That was already a meaningful change, but it did not prove much beyond this particular page.</p><p>So I tried five more pages from <em>Sony&#335;n</em>, using prose and dialogue rather than word games. These pages contained a wider range of old Korean forms: arae-a in forms such as &#4370;&#4510; and &#4354;&#4510;&#4523;, several &#12613;-cluster groups such as &#12666;, &#12668;, and &#12669;, and even the rarer triple cluster &#12660;. This was the real test. The question was no longer whether a model could preserve two repeated forms on one page, but whether it could keep a broader old-orthography repertoire across different pages.</p><p>The current models can preserve old Korean forms, but not with equal reliability. Preservation is still something to check, not something to assume. In practice I treated an old form as trustworthy only when different models produced it, because a single model can also hallucinate plausible old-looking Korean.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sTf9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sTf9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 424w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 848w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 1272w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sTf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png" width="880" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:880,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28408,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!sTf9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 424w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 848w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 1272w, https://substackcdn.com/image/fetch/$s_!sTf9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2907f57-8f12-4389-9858-cee2c0eea3dc_880x460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The dedicated OCR engines were not trained on any form of old-Korean and not surprising that they performed the poorest. <a href="https://mistral.ai/news/ocr-4/">Mistral OCR</a>, <a href="https://github.com/deepseek-ai/DeepSeek-OCR">DeepSeek-OCR</a>, <a href="https://github.com/baidu/Unlimited-OCR">Baidu&#8217;s Unlimited-OCR</a>, <a href="https://github.com/zai-org/GLM-OCR">GLM-OCR</a>, all read the page and then modernized the old letters or fell apart, one of them looping a few character lines a few hundred times, another reading the Hangul as Chinese and drifting into a gibberish. Qwen3-VL was different. It read the Korean fluently, but silently regularized every older form into modern Korean. If your goal is simply to recover the text, that is a respectable result. If you care about preserving historical orthography, it is not.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ysBP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ysBP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 424w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 848w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 1272w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ysBP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png" width="1456" height="1399" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1399,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:553025,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ysBP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 424w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 848w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 1272w, https://substackcdn.com/image/fetch/$s_!ysBP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f0b5486-551b-4fdc-8515-8595a5e0ab3b_2271x2182.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Going back further, to 1896</h2><p>After <em>Sony&#335;n</em>, I wanted to know how the models handled even earlier Korean print. As a final test, I tried the front page of the <em>Tongnip Sinmun</em> (&#46021;&#47549;&#49888;&#47928;, <em>The Independent</em>) from 1896, the first privately published Korean newspaper, set entirely in Hangul in the full old orthography with arae-a in almost every word. The only scan I could find is small, a single 960-pixel page, which makes it a hard test on top of an old one. For once there was something solid to check it against. The Korean Wikisource was manually transcribed, so this is the one page in the whole exercise I could measure against ground truth rather than judge by eye. It should also be noted that this newspaper image is denser and has a more difficult layout then the previous magazine had.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qQK7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qQK7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qQK7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg" width="960" height="1307" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1307,&quot;width&quot;:960,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:525341,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qQK7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qQK7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bc14bf-d20b-4878-b554-5937979a0206_960x1307.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most models struggled. Claude Opus 4.8 transcribed only about a sixth of the page, even reversing the masthead. Qwen3-VL reached roughly two thirds, but regularized the spelling throughout. Both ended up at around ninety percent character error against the Wikisource transcription.</p><p>Gemini 3.1 Pro was the clear exception. It read essentially the entire page, from the masthead through the advertisements and price table to the end of the editorial, about ninety-six percent of the page by length. Compared against the Wikisource transcription, it reached roughly seventeen percent character error while preserving 176 of the page&#8217;s 182 arae-a. On a 960-pixel scan of an 1896 newspaper, that amounts to a complete transcription of a full old-orthography page.</p><p>GPT-5.5 failed for a different reason. It did not misread the page so much as fail to read it at all. <em>Tongnip Sinmun</em> is one of the most famous newspapers in Korean history, and its first issue has been reproduced countless times. When shown the thirteenth issue, GPT-5.5 instead reproduced the first: its date, issue number, and opening lines, at both image resolutions. Against the actual page, the result is about eighty percent character error. It is a useful reminder that, for well-known historical documents, a model can sometimes retrieve what it already knows instead of reading what is actually in front of it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IRCS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IRCS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 424w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 848w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 1272w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IRCS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png" width="1456" height="1567" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1567,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1571153,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IRCS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 424w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 848w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 1272w, https://substackcdn.com/image/fetch/$s_!IRCS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F456ef56c-adad-49b9-bf91-c5366ae3718f_2399x2582.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>Why this matters to someone working with the sources</h2><p>The interesting part is not the benchmark. It is what happens to the historical record.</p><p>When the National Institute of Korean History transcribes historical print, it modernizes the spelling. I checked the roughly six hundred thousand colonial-era records they have published, and almost none preserve the older letters. For anyone searching the archive, the original orthography has effectively disappeared.</p><p>That makes preservation more than a technical curiosity. A model that keeps the old letters is not simply producing another transcription. It is recovering information that was lost during digitization. For anyone working with pre-1933 Korean print, this may be of importance.</p><p>The preservation results only matter if the models can also read the page accurately. On that front, the picture is encouraging. Against four pages of the 1921 magazine <em>Kaeby&#335;k</em>, where a modern transcription exists, the best current models matched roughly ninety-seven percent of the printed characters directly from the low-resolution scans. That measures reading rather than preservation, since the 1921 transcription already modernizes the spelling. Taken together, though, the results point in the same direction. Current models can read these pages, and on sufficiently old material, the best of them also preserve what makes them historically distinctive.</p><h2>A note on scale</h2><p>For a single page, I would use a frontier vision-language model every time.</p><p>My own project is different. It involves millions of character crops. Running a frontier model over that volume would incur substantial API costs.</p><p>Instead, I built a small voting system with more traditional OCR engines. EasyOCR and a lightweight classifier trained on historical Korean type recognize the Old-Korean, PaddleOCR reads the Hanja, and a simple voting rule chooses between them. The entire corpus finished in about a week on hardware we already owned.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1Ve-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1Ve-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 424w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 848w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 1272w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1Ve-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png" width="1456" height="1361" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1361,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:323713,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1Ve-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 424w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 848w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 1272w, https://substackcdn.com/image/fetch/$s_!1Ve-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcca54901-d06c-4486-be1e-318ec21e2b74_1832x1713.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The ensemble has one advantage over a frontier model besides cost. It is explicit about uncertainty. When its engines disagree, it leaves the character blank instead of guessing. A frontier model often does the opposite. It produces a perfectly plausible modern character, and unless you compare it with the original page you may never realize the historical form has disappeared. A blank is an error you can revisit. Silent modernization is much harder to detect.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wSGI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wSGI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 424w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 848w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 1272w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wSGI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png" width="1456" height="2312" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2312,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:939561,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203420919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wSGI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 424w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 848w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 1272w, https://substackcdn.com/image/fetch/$s_!wSGI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c98e488-4dc8-4155-891c-a3daa818873f_1514x2404.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What this adds up to</h2><p>For anyone working with pre-1933 Korean print, the practical picture has changed. Current vision-language models can do more than transcribe these pages. The strongest of them can preserve historical orthography that is absent from most digital editions. It can also help one quickly digitize small (or large) amount of works. However for very large datasets cost can quickly skyrocket. It might be worth distilling some of this knowledge into a smaller, more specific model to save on API costs. </p><p>That does not mean OCR has been fully solved. The results are model-dependent, page-dependent, and, as far as I can tell, language-dependent. But for this particular corner of the problem, the gap between what was possible a few years ago and what is possible today is surprisingly large.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Chinese-Korean-Japanese </p><p></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Kim, Michael. &#8220;The Han&#8217;g&#365;l Crisis and Language Standardization: Clashing Orthographic Identities and the Politics of Cultural Construction.&#8221; Journal of Korean Studies 22, no. 1 (March 2017): 5&#8211;31. https://doi.org/10.1215/21581665-4153412.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[A council of models]]></title><description><![CDATA[Exploring how we can use LLMs like a team for empirical research.]]></description><link>https://www.pixelsandpatterns.org/p/a-council-of-models</link><guid isPermaLink="false">https://www.pixelsandpatterns.org/p/a-council-of-models</guid><dc:creator><![CDATA[Steven Denney]]></dc:creator><pubDate>Fri, 26 Jun 2026 13:39:34 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fb5cf221-5ac9-4a62-9d0c-4b536bbc932e_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IoIu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IoIu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IoIu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2086548,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203542307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IoIu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!IoIu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c9b982-5166-43a3-8eac-74b5e61d8e2d_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>This is an extension of part of the <a href="https://scdenney.github.io/assets/slides/from-pixels-to-patterns/#1">presentation</a> on using LLMs in research pipelines I gave for <a href="https://machinecollaborators.org/">Machine Collaborators</a>, series of talks organized by Charles Crabtree (Monash University) where researchers walk through how AI is being used in their work.</em></p><p>OpenRouter recently announced a product called <a href="https://openrouter.ai/blog/announcements/fusion-beats-frontier/">Fusion</a>, which routes a single question through a panel of cheaper models at once and reports that such a panel can come within a point of a frontier model's score at about half the cost. The marketing is new, but the idea a bit less so. Pitting several models against one another has become a standard way to use language models for finding optimal ideas and outcomes. For empirical research, what use might it have?</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><p>I have used a similar setup in two of my own studies. In one, six models read roughly four thousand open-ended survey answers and sorted each into a code, helping me determine what respondents were actually thinking when they navigated a survey experiment. In the other, four models read sixty-seven Korean history textbooks and, using a term extraction technique I created, determined which words best reflected national identity. Sorting answers into fixed codes and discovering which words matter are different problems, but the instrument is the same. It is a panel of open-weight language models, treated as a team of independent coders, voting.</p><p>In this piece I reflect on how it works, why I run it mostly on open-weight models, how it relates to the &#8220;councils&#8221; that Fusion and others have started to build, and what it does and does not support you, as a researcher, in doing. The short version is that it works well, that for coding tasks it holds up against the human labor it replaces, and that its central value is scale rather than certainty.</p><h2>The method: open models as a panel of coders</h2><p>Take a judgment a trained research assistant would make, such as reading a piece of text and deciding what it is about, and give it to several models at once, each with the same codebook and the same instructions. Each model receives the same text, the same instructions, and the same codebook, then returns a single labeled decision. A response receives a classification, while the number of models supporting that decision is recorded as a measure of agreement.</p><p>Two features of this design matter for the steps that follow. First, I treat each model as a coder. I ask it to assign a label, much as I would ask a trained research assistant, rather than relying on its probability distribution or confidence estimates. (More on this below.) I ask it for a labeled decision, the way I would ask a research assistant, and I read the spread of decisions across the panel the way content analysts read inter-coder reliability, with chance-corrected statistics such as Cohen&#8217;s and Fleiss&#8217; kappa. Second, the models vote independently. They run at temperature zero (i.e., no &#8220;creativity&#8221; permitted), they do not see one another&#8217;s answers, and there is no deliberation. Any sense in which the models are played off against one another happens afterward when I line the votes up and read where they part ways.</p><p>I run the panel mostly on open-weight models. A proprietary model behind an API can change between versions, and its outputs cannot be reproduced exactly by another researcher. Open-weight models, run locally at a fixed temperature and seed, return the same answers when re-run in the same environment, which is what reproducible measurement requires. In the classification study I still included one proprietary model, <a href="https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/">GPT-4o-mini</a>, as a convenient reference coder, because it is inexpensive, fast, and unusually capable on the Korean and Traditional Chinese text that those projects involved. The term-extraction panel below is fully open-weight. But GPT-4o-mini is one voice on a panel that is otherwise open and re-runnable, not the authority on it.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B_01!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B_01!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 424w, https://substackcdn.com/image/fetch/$s_!B_01!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 848w, https://substackcdn.com/image/fetch/$s_!B_01!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 1272w, https://substackcdn.com/image/fetch/$s_!B_01!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B_01!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png" width="1456" height="852" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:852,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:244040,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203542307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B_01!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 424w, https://substackcdn.com/image/fetch/$s_!B_01!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 848w, https://substackcdn.com/image/fetch/$s_!B_01!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 1272w, https://substackcdn.com/image/fetch/$s_!B_01!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2436f8a-977d-48b9-88e8-763224d4941c_1600x936.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>An example of the method. One open-text response from a survey is read by six models that vote independently. The majority becomes the code, and the size of the majority is recorded as a reliability signal.</em></p><h2>Three ways to convene a council</h2><p>The instinct to put models in concert has been built in several distinct ways, and they are worth separating, because they are not after the same thing.</p><p><a href="https://openrouter.ai/blog/announcements/fusion-beats-frontier/">Fusion</a> is the production version. A panel of models answers in parallel, a judge model reads every answer and writes a structured account of where they agree, where they contradict, and what they all missed, and the calling model writes the final answer grounded in that account. What Fusion optimizes is the quality of one answer against its cost, which is why its own documentation says it is not a drop-in replacement and is worth the expense only when a question warrants several perspectives. The models often arrive at different conclusions, but those differences are resolved internally rather than exposed to the user. </p><p>Sakana AI&#8217;s <a href="https://sakana.ai/fugu-release/">Fugu</a> uses the same basic approach. Rather than sending every query to a fixed panel of models in parallel, it uses a trained orchestration model to decide whether and when to delegate, select and coordinate models from an agent pool, verify their work, and synthesize the results into a single answer. The system is exposed through one model API, so the complexity of multi-model coordination never reaches your code.</p><p>Two other versions come from computer and political science. Andrej Karpathy's <a href="https://github.com/karpathy/llm-council">llm-council</a> has several models answer, then rank one another with their identities hidden, before a chairman model writes a final answer. Andrew Hall's <a href="https://github.com/andybhall/llm-council-governance">llm-council-governance</a> extends that logic into an experiment on the decision rule itself, comparing seven governance procedures, from simple majority voting to deliberate-then-vote, on problems with verifiable answers. Deliberation followed by a vote performed best, at 80.9 percent against 71.7 for the strongest single model. There is also an academic <a href="https://arxiv.org/abs/2406.08598">Language Model Council</a> in which twenty models write tests, answer them, and grade one another to produce rankings that track human judgment more closely than any single model does.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vWqz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vWqz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 424w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 848w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vWqz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png" width="1456" height="809" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:809,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:292561,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203542307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vWqz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 424w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 848w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!vWqz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30b4ed4e-4598-4fe0-bfb2-208473ab0ad6_2160x1200.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Three ways to wire a panel of models. All draw on the same lineage. They differ in what they do with the disagreement.</em></p><p>The resemblance is close enough that the model lineups overlap. Hall's council runs Qwen, Llama, Gemma, and Mistral. The classification panel below runs Qwen, Llama, Gemma, and three others. What separates my use from the rest is the last step. Fusion and the councils collapse the panel into a single output, an answer or a decision, and discard the spread behind it. Hall's experiment even finds that deliberation before the vote produces the best single answer, which is the right goal when you want one answer and the wrong one when the spread is the thing you are trying to measure. So I keep the spread, because in a measurement task the agreement, or the lack of it, is itself a quantity of interest. It helps to think of each model as a coder rather than a black box. Each one assigns a label, the panel may agree or divide, and each carries some uncertainty of its own. Those are the qualities I want to read, not average away.</p><h2>Application one: classification</h2><p>Take classification first. My thinking on this grew out of my working paper, <a href="https://github.com/scdenney/what-were-they-thinking">"What Were They Thinking?"</a>, which asks how to validate survey constructs in conjoint designs using the open-text data usually collected only for manipulation checks. A conjoint experiment tells you what respondents chose, but rarely whether they reasoned about the choice the way the design assumes. To find out, I embedded an open-ended question after the conjoint task ("Why did you choose this person?") and coded the answers into a small set of categories. The corpus is about four thousand short answers in Korean and Traditional Chinese, from naturalization experiments in South Korea and Taiwan, where respondents weighed <a href="https://github.com/scdenney/cues-east-asia">who deserves to be prioritized for citizenship</a>. The codebook the models worked from had five codes: civic commitment, functional integration, economic focus, identity concern, and a none-of-the-above residual.</p><p>Six models coded every answer: GPT-4o-mini together with five open-weight models run locally through <a href="https://ollama.com/">Ollama</a>, spanning <a href="https://qwenlm.github.io/blog/qwen2.5/">Qwen 2.5</a> at 3B, 32B, and 72B, <a href="https://ai.meta.com/blog/meta-llama-3-1/">Llama 3.1</a> 8B, and <a href="https://ai.google.dev/gemma">Gemma 3</a> 12B. A response is given a classification when at least four of the six agree. Across the corpus the panel&#8217;s chance-corrected agreement is a Fleiss&#8217; kappa of about 0.74, which content analysts would call substantial. As more of the six models converge on a label, each model is also individually more confident in it, and the unanimous cases are, likely, the ones a researcher could reasonably trust. The split cases, where the panel divides three against three, are the cases worth handing to a human reader or dealing with in another way.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JaTI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JaTI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 424w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 848w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JaTI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png" width="1456" height="849" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:849,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:100849,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203542307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JaTI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 424w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 848w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!JaTI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78850d85-a963-4166-82c1-41df79a374c6_1920x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>As more of the six models agree on a code, each model is also individually more confident in it. Most answers sit at high agreement. The split cases are the ones worth reviewing manually.</em></p><p>Two things make this more than a cost-saving trick. The first is scale. Hand-coding four thousand multilingual answers to a publishable standard, with two trained coders and an adjudication protocol, is a substantial project on its own, and it does not scale to the next survey or the one after that. A panel of models codes the full corpus in an afternoon and recodes it with ease, assuming a replicable setup, if the codebook is updated. </p><p>The second is quality. I do not treat the models as a cheap approximation of a human coder so much as a fast one that reasons about a short, context-dependent answer in much the way a careful human reader would. Where I had human-coded ground truth and the panel disagreed with it, I read the disputed cases myself, and more often than not I judged the model's label the better one.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> This is consistent with <a href="https://osf.io/preprints/socarxiv/zr5vf_v1">Ryan Briggs and colleagues</a>, who used language models to extract validated data from roughly one hundred thousand political science articles, putting model coding to work at a scale hand-coding could not reach.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>Not every disagreement is the model's mistake to absorb, though. Some are a sign that the codebook is unclear, and those are the ones to chase, because the codebook is the part you actually tune. When I first checked the panel against human-coded ground truth, the agreement was only fair. Reading the misses in a batch, rather than one at a time, showed the models were applying two of my categories the way I had written them and not the way I meant them: one defined too narrowly, the catch-all residual left too broad. I rewrote those definitions, ran the panel again, and looked again, and over two rounds the agreement with the human codes rose from a fair kappa near 0.40 to a substantial one near 0.67. That gain was larger than anything I got from swapping in a bigger model. The model is the coder. The codebook is its instructions, and calibrating the instructions against a few hundred of your own annotations is where most of the accuracy is won.</p><h2>Application two: finding the words that matter</h2><p>The second use case approaches the use of multiple LLMs a bit differently. In <a href="https://scdenney.net/research/">"Constructing the Nation,"</a> written with <a href="https://aronvandepol.com/">Aron van de Pol</a>, I study how South Korean history textbooks portray national identity across the postwar decades. This time there is no codebook to apply. The categories themselves are the quantity of interest. We set out to find which words carry the weight of national identity in the corpus, and how their use changes over time.</p><p>The standard tool for extracting terms, concepts, or words from a corpus is topic modeling. The standard topic model is <a href="https://www.jmlr.org/papers/v3/blei03a.html">Latent Dirichlet Allocation</a>. LDA treats each document as a bag of words and groups terms that co-occur. It is fast and unsupervised, and it is indifferent to context. A word that means one thing in a sentence about colonial resistance and another in a sentence about economic development is the same token to LDA. For a corpus of long, discursive textbook chapters, where the meaning of a term such as <em>minjok</em> (nation, or ethnos) gets its meaning from the sentences around it, the lack of context being taken into account is a real limitation.</p><p>An LLM approach handles it differently. Four large models, each from a different country's lab, read the corpus in long passages and propose the identity-related terms they find: <a href="https://huggingface.co/LGAI-EXAONE">EXAONE</a> from Korea, <a href="https://cohere.com/research/aya">Aya Expanse</a> from Canada, Qwen from China, and Gemma from the United States. A term enters the final set only when at least three of the four propose it, after which each candidate has to clear a corpus-frequency and semantic-similarity check. The models were chosen from different training traditions on purpose, so that agreement across them would mean more than agreement among near-copies. The result is a set of nine defining terms.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x1te!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x1te!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 424w, https://substackcdn.com/image/fetch/$s_!x1te!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 848w, https://substackcdn.com/image/fetch/$s_!x1te!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!x1te!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x1te!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png" width="1456" height="849" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:849,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:91605,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://pixelsandpatterns.substack.com/i/203542307?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!x1te!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 424w, https://substackcdn.com/image/fetch/$s_!x1te!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 848w, https://substackcdn.com/image/fetch/$s_!x1te!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!x1te!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe1c61d0-87a8-4bf9-84a1-68c75505d310_1920x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Of the nine terms the four-model consensus defines, an embedding-based topic model identifies all nine, while classic LDA identifies five. Context-aware methods find terms that bag-of-words co-occurrence misses.</em></p><p>As a check, we ran an embedding-based topic model, <a href="https://maartengr.github.io/BERTopic/">BERTopic</a>, which reads words in context. It identified all nine, with far fewer edge cases than the bag-of-words baseline. BERTopic does not read a textbook whole, though. Its embedding model takes in short passages at a time, because its context window holds only so many tokens.</p><p>An LLM faces the same limit, only a much larger one. It cannot take in a whole book at once either, since the text would exceed its context window and be rejected or truncated, but its window is wide enough to read a term against far more of the surrounding passage, and you can window across a long document in overlapping batches. That extra room is where the real difference shows up.</p><p>Classic LDA, for what it is worth, identifies only five terms. The panel of LLMs does what topic modeling is meant to do: name the vocabulary that organizes a corpus. But it reads each term in the context of the passage around it, the way a human reader of these textbooks would, and it identifies terms that word co-occurrence alone misses. For a corpus of long documents, that is the difference between a list of frequent words and a list of meaningful ones.</p><h2>What LLM voting does and does not measure</h2><p>There are, of course, limits. Cross-model agreement is a measure of reliability, not of validity, and it is not a calibrated estimate of uncertainty. The panel can agree because the answer is clear, and it can agree because the models share training data and therefore share blind spots. The Condorcet jury theorem, which is the formal reason a majority of independent voters grows more accurate as the panel grows, assumes the voters are independent, and language models are not. Their errors are correlated, so a panel of nine carries far less independent information than nine human coders would.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> Agreement should therefore be reported as inter-coder consistency among non-independent coders, rather than accuracy or probability.</p><p>This limit is notable, but it is not, in my experience, the main thing. On the one task where I had adjudicated human codes to check against, which came from a separate corpus on North Korean entrepreneurs rather than the citizenship answers above, the models agreed with the human coders at a substantial level, short of the agreement between the humans themselves, and in the cases where they differed I usually sided with the model. The limit imposes modest discipline. Keep some human-coded ground truth in the project, if you can, even a few hundred cases, so that the panel can be checked rather than trusted on faith. Where a within-model confidence is needed, read it from the model's own token probabilities, which come from its predictive distribution, rather than from the agreement count or a model's self-reported confidence.</p><p>That last point is worth dwelling on, because if a model is going to stand in for a coder, its uncertainty has to be handled like a coder&#8217;s. A careful human coder is not equally certain of every judgment, and neither is a language model. Even with a fixed codebook and temperature set to zero, the model assigns different degrees of confidence to different classifications. That confidence is reflected in the <a href="https://cookbook.openai.com/examples/using_logprobs">log-probability</a> it assigns to the label it ultimately selects, a quantity derived from the model&#8217;s predictive distribution at the moment of the decision. By contrast, a HIGH or LOW confidence tag the model produces if asked is simply another generated response, not a direct measurement of its uncertainty.</p><p>I therefore treat the log-probability as an indicator of uncertainty. It records how much probability the model assigned to the label it ultimately selected, according to its predictive distribution. A response that receives a label but a relatively low log-probability is one that merits closer inspection. It captures something the panel does not. The log-probability reflects how certain each model is about its own decision. Agreement across the panel reflects whether different models reach the same decision. The two measures complement one another rather than serving the same purpose.</p><h2>Is it worth it</h2><p>It is, if the claim stays within the evidence. A panel of open-weight models is a practical research instrument, and the two applications here show its range. For classification it does the work of a team of coders at a scale no team could match, and codes nearly as well, a little below expert human agreement. For term discovery it does the work of a topic model while reading each term in context, which lets it identify defining vocabulary that bag-of-words methods miss. In both cases the design that makes it trustworthy is the same. Use several models rather than one, choose them to be as different from one another as possible, run them on open weights so the result can be reproduced, and treat a split vote as a flag for human attention rather than as noise to be averaged away.</p><p>A panel of LLMs can outperform a single model. But it does not follow that high agreement means the panel is correct. High agreement tells us that these particular models concur, not that they have reached the truth. The two coincide only when the panel is both accurate and sufficiently diverse. Build the panel and pay attention to where it disagrees. But the human stays very much in the loop here, refining the codebook before the run and adjudicating the cases the panel cannot resolve afterward. The panel supplies scale, consistency, and a record of disagreement. The researcher still has to decide what the categories mean, whether the instrument is valid, and when the models and their output should not be trusted.</p><h2>Using Claude Code</h2><p>If you want to run something like this yourself, the steps are codified as a set of open-source <a href="https://github.com/scdenney/open-science-skills">Claude Code skills</a> I maintain and reach for when working with agents, so each run stays reproducible. The ones closest to this post are <em>model-council-voting</em>, for running a panel and reading its agreement; <em>text-classification</em>, for designing and validating a codebook; and <em>llm-calibration-logprobs</em>, for reading per-decision confidence from token log-probabilities.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.pixelsandpatterns.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.pixelsandpatterns.org/subscribe?"><span>Subscribe now</span></a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>The full classification panel is GPT-4o-mini plus five open-weight models run locally via <a href="https://ollama.com/">Ollama</a> at temperature zero with a fixed seed: <a href="https://qwenlm.github.io/blog/qwen2.5/">Qwen 2.5</a> (3B, 32B, 72B), <a href="https://ai.meta.com/blog/meta-llama-3-1/">Llama 3.1</a> (8B), and <a href="https://ai.google.dev/gemma">Gemma 3</a> (12B). The case for running open weights rather than depending on a proprietary API is reproducibility: API models drift between versions and cannot be rerun deterministically (Barrie et al. 2025). GPT-4o-mini earns its place as a reference coder on grounds of cost and strong multilingual performance, not authority, and agreement with it rises monotonically with open-model size (Qwen 2.5: 0.73 at 3B, 0.83 at 32B, 0.86 at 72B), which is a useful sanity check on the smaller models.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>The check comes from a separate corpus coding why North Korean entrepreneurs make the choices they do, where two trained humans coded every response. The two humans agree at kappa 0.88. Each model agrees with the humans at about 0.61 to 0.67, roughly what one would expect of a competent third coder. The models are also overconfident against human codes when their token probabilities are read as calibrated estimates (expected calibration error of 0.15 to 0.26, against 0.01 to 0.10 when scored against the panel&#8217;s own majority), which is the technical form of the point that agreement is not the same as accuracy.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Briggs, Mellon, and Arel-Bundock, <a href="https://osf.io/preprints/socarxiv/zr5vf_v1">&#8220;It must be very hard to publish null results&#8221;</a> (working paper). They code a large corpus of articles for whether results are null, using LLMs validated against human coding, and document a steep gap between the share of published abstracts reporting non-null results and the share reporting nulls.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>A recent study, <a href="https://arxiv.org/abs/2605.29800">&#8220;Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels&#8221;</a> (Kohli 2026), estimates that a panel of nine frontier models from seven families carries only about two independent votes&#8217; worth of information, because the models tend to be wrong on the same items. The remedy is real diversity among the voters, which is the reason the term-discovery panel draws its four models from four different national labs.</p></div></div>]]></content:encoded></item></channel></rss>