<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Local A.I, Applied. blog by ThirdShift R&D]]></title><description><![CDATA[Getting local AI across the gap to something a regular business can run safely, on hardware it owns. Honest, reproducible eval: tests, methods, numbers with error bounds — and the failures. Build or deploy? Follow along.]]></description><link>https://research.thirdshiftrd.com</link><image><url>https://research.thirdshiftrd.com/img/substack.png</url><title>Local A.I, Applied. blog by ThirdShift R&amp;D</title><link>https://research.thirdshiftrd.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 17 Sep 2026 12:59:47 GMT</lastBuildDate><atom:link href="https://research.thirdshiftrd.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Local A.I. Applied]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[3rdshiftrnd@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[3rdshiftrnd@substack.com]]></itunes:email><itunes:name><![CDATA[Local A.I. Applied]]></itunes:name></itunes:owner><itunes:author><![CDATA[Local A.I. Applied]]></itunes:author><googleplay:owner><![CDATA[3rdshiftrnd@substack.com]]></googleplay:owner><googleplay:email><![CDATA[3rdshiftrnd@substack.com]]></googleplay:email><googleplay:author><![CDATA[Local A.I. Applied]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[04 - What does local A.I. actually cost to run?]]></title><description><![CDATA[On our hardware, the answer is less about the token price than the hours you leave the box waiting]]></description><link>https://research.thirdshiftrd.com/p/04-what-does-local-ai-actually-cost</link><guid isPermaLink="false">https://research.thirdshiftrd.com/p/04-what-does-local-ai-actually-cost</guid><dc:creator><![CDATA[Local A.I. Applied]]></dc:creator><pubDate>Tue, 18 Aug 2026 08:16:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/da71819b-9287-415d-b17e-d92285886e87_1200x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>[drafted by project agent, edited by me, by hand]</p><p>Ask what local AI costs and the conversation usually jumps straight to the hardware bill. The other side answers with an API price per million tokens. Both numbers are real. Neither tells you what it costs to keep a local box available, or what one more answer adds to the meter.</p><p>So we measured ours.</p><p>The short version is this:</p><blockquote><p><em><strong><span>On our hardware, a typical local answer costs about $0.00003 in marginal electricity at the measured operating point &#8212; roughly 0.003&#162;. Keeping the box powered and ready is the larger number: about $10 a month running 24/7 at our all-in reference rate.</span></strong></em></p></blockquote><p>Those are not universal prices. They are numbers for one box, one model, one serving mode, one workload, and one electricity rate. The useful part is the shape of the cost, and the method is simple enough to repeat.</p><h3><strong>The box and the meter</strong></h3><p>The measured machine is an Intel Arc Pro B60 attached to a UM890 Pro host through an OCuLink dock. We measured both halves at the wall with two smart plugs: one for the host and one for the card, dock, and dedicated power supply.</p><p>The local model for the benchmark point was a Qwen3.6-35B-A3B Q4_K_M served with llama.cpp, single stream. The wall-meter benchmark used a Tapo P115 with stated accuracy of approximately &#177;2%.</p><p>The monthly number came from a different, longer measurement: both plugs ran for about 25.3 days with no recorded off-time.</p><p>That distinction matters. A short benchmark tells you about the energy used while answering. A month-long plug reading tells you what the machine costs when it is part of the room, including the time it spends waiting.</p><h3><strong>What the month-long reading showed</strong></h3><ul><li><p>Measured energy, about 25.3 days - <strong>35.99kWh</strong> </p></li><li><p>Average draw over that period - <strong>59.2W</strong></p></li><li><p><strong>Projected 31-day</strong> energy  -<strong> </strong><em>about</em><strong> 44 kWh</strong></p></li><li><p>All-in reference rate used for the conversion - <strong>$0.2208/kWh</strong></p></li><li><p><strong>Projected monthly</strong> electricity cost - <em>about</em><strong>  $9.72</strong></p></li><li><p>Average <strong>cost per hour</strong>, running 24/7 - <em>about</em><strong> 1.3&#162;</strong></p></li></ul><p>The plain-English version is <strong><span>under $10 a month to run 24/7</span></strong>, at this rate and for this measured rig. It works out to about <strong><span>29&#162; a day</span></strong>.</p><p>The average draw is close to the measured idle floor because the box is not generating continuously. It spends much of its life powered on and waiting. That is the idle tax, and it is the number a marginal-only calculation leaves out.</p><h3><strong>What one more answer adds</strong></h3><p>The marginal answer number comes from the locked wall-meter operating point, not from dividing the monthly bill by an arbitrary number of prompts.</p><p>At the balanced chat operating point, the whole box measured about <strong><span>799 output tokens per watt-hour</span></strong> in the lock-grade run. A served 8K single-stream operating point measured about <strong><span>671 tokens per watt-hour</span></strong>. The energy methods agreed to better than 1% on the benchmark windows; the absolute meter accuracy remains the relevant &#177;2% bound.</p><p>Using the recorded benchmark conversion, the preliminary marginal estimate is:</p><blockquote><p><em><strong><span>about $0.00003 per answer at $0.17/kWh &#8212; approximately 0.003&#162;.</span></strong></em></p></blockquote><p>Using the higher all-in rate from the measured household bill, $0.2208/kWh, the same operating-point estimate scales to roughly <strong><span>$0.00004 per answer</span></strong>. We are rounding this as an order-of-magnitude field number, not claiming that every answer has the same length or energy profile.</p><p>A short answer, a long answer, a cold uncached prompt, and a decode-heavy chat do not cost the same. The benchmark itself showed three regimes:</p><ul><li><p><strong><span>READ:</span></strong> about 297 tok/Wh, cold and uncached, with an 8,008-token input prefilled per answer.</p></li><li><p><strong><span>CHAT:</span></strong> about 799 tok/Wh, the balanced middle case.</p></li><li><p><strong><span>WRITE:</span></strong> about 1,020 tok/Wh, a decode-dominated ceiling.</p></li></ul><p>That is why the honest sentence is <strong><span>&#8220;on our hardware, about X,&#8221;</span></strong> not <strong><span>&#8220;local AI always costs X.&#8221;</span></strong></p><h3><strong>The part that changes the answer: whether the box is already on</strong></h3><p>The marginal number is the cost of doing one more job while the machine is already running. It does not pay for the hardware, the dock, the host, or the hours spent waiting.</p><p>If the box is already powered for other work, the extra electricity for one answer can be close to the marginal number. If the only reason to keep it running is one occasional question, the idle draw becomes the dominant cost.</p><p>That gives local AI two different economic stories:</p><ol><li><p><strong><span>High use:</span></strong> the fixed cost of keeping the box on is spread across more answers. The marginal electricity is small.</p></li><li><p><strong><span>Low use:</span></strong> the idle tax can matter more than the answers. A box that sits waiting all month is not &#8220;free&#8221; just because each generated answer uses a fraction of a cent.</p></li></ol><p>Neither story is the whole answer. You need both.</p><h3><strong>What this does not include</strong></h3><p>These figures are an electricity floor plus a measured operating-cost view. They do not include:</p><ul><li><p>the purchase price or depreciation of the hardware;</p></li><li><p>maintenance, replacement parts, or downtime;</p></li><li><p>the operator&#8217;s time;</p></li><li><p>networking, storage, or other household infrastructure;</p></li><li><p>the quality of the answer compared with a cloud service;</p></li><li><p>a universal electricity rate.</p></li></ul><p>The measured monthly figure includes the actual all-in rate used for our internal conversion. Readers should substitute their own rate. The published method should carry the watt-hours and assumptions without exposing a location-specific utility bill.</p><h3><strong>So, what does local A.I. cost?</strong></h3><p>On this box, at this operating point:</p><ul><li><p><strong><span>One additional answer:</span></strong> about <strong><span>$0.00003&#8211;$0.00004</span></strong> in electricity, or roughly <strong><span>0.003&#8211;0.004&#162;</span></strong>, depending on the reference rate used.</p></li><li><p><strong><span>Keeping the complete rig powered 24/7:</span></strong> about <strong><span>$10 a month</span></strong> at the measured all-in rate.</p></li><li><p><strong><span>The practical warning:</span></strong> at low usage, idle power matters more than the cost of the individual answer.</p></li></ul><p>That is the number we can stand behind: not a promise for every local-AI build, but a measured answer for ours, with enough method attached that someone else can check the shape against their own wall meter.</p><p>The cost is small. The assumptions are not.</p><p></p><h5>The method</h5><ul><li><p><strong><sub><span>Whole rig:</span></sub></strong><sub> UM890 Pro host plus Intel Arc Pro B60, OCuLink dock, and dedicated PSU.</sub></p></li><li><p><strong><sub><span>Monthly telemetry:</span></sub></strong><sub> two smart plugs, approximately 25.3 days, 35.99 kWh total, 59.2 W average.</sub></p></li><li><p><strong><sub><span>Projected monthly conversion:</span></sub></strong><sub> approximately 44 kWh &#215; $0.2208/kWh = approximately $9.72.</sub></p></li><li><p><strong><sub><span>Benchmark:</span></sub></strong><sub> Tapo P115, approximately &#177;2%; Qwen3.6-35B-A3B Q4_K_M; llama.cpp; single stream; temperature 0; lock-grade run dated 2026-07-15.</sub></p></li><li><p><sub>The marginal per-answer figure is preliminary for this article&#8217;s framing. It is not a universal tariff or a fully loaded total cost of ownership.</sub></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://research.thirdshiftrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Local A.I, Applied. blog by ThirdShift R&amp;D! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[03 - Recency isn't authority]]></title><description><![CDATA[Why AI memory needs provenance, not just lifecycle states]]></description><link>https://research.thirdshiftrd.com/p/recency-isnt-authority</link><guid isPermaLink="false">https://research.thirdshiftrd.com/p/recency-isnt-authority</guid><dc:creator><![CDATA[Local A.I. Applied]]></dc:creator><pubDate>Thu, 30 Jul 2026 15:53:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f8a1a5d4-2d54-4f56-91a7-1376c3660ff1_1456x819.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>(drafted by project agent, edited by hand, by me)</p><h3>Why we can&#8217;t rely on Recency</h3><p>A business&#8217;s own documents are not a fixed truth. They pile up. The safety procedure written in January gets revised in March and rewritten in June. The old one doesn&#8217;t vanish when the new one lands&#8230; it&#8217;s still in the drawer, still on the server, still perfectly readable. All three are real. All three were correct on the day they were written.</p><p>Feed that drawer to an AI and you&#8217;ve handed it a problem no amount of better retrieval solves. Everyone&#8217;s trying to fix memory by improving the fetch; better RAG, better embeddings, bigger context windows. But when three versions of the same procedure all legitimately match the question, retrieval isn&#8217;t the bottleneck. Retrieval does its job perfectly, it hands you all three. Now what?</p><h4>Why Provenance is necessary</h4><p>The tempting answer is recency. Take the newest one. It&#8217;s clean, it&#8217;s cheap, and it&#8217;s wrong often enough to hurt. The newest document isn&#8217;t automatically the governing one. A later draft isn&#8217;t a ratified policy. A memo that mentions the procedure doesn&#8217;t replace it. An addendum changes one clause, not the whole document. &#8220;Most recent&#8221; is a guess that fails where the stakes are highest, the moment the answer actually depends on which version rules.</p><p>Lifecycle states are a real step up from a flat pile. Marking a memory Current, Superseded, or Stale beats treating everything as equally true forever. But states only move the question, they don&#8217;t answer it. Superseded <em>by what?</em> Current <em>on whose authority?</em> A label with nothing underneath it is just a guess wearing a badge. You&#8217;ve renamed the problem, not solved it.</p><h4>Applying to AI</h4><p>None of this is new, either. Anyone who&#8217;s built a data warehouse has seen it. Bi-temporal modeling and slowly-changing dimensions solved the <em>shape</em> of this decades ago, and event sourcing keeps the history searchable. What&#8217;s new is needing it one layer up, in the memory an AI reads from and answers out of. Old discipline, less forgiving place.</p><p>The part that actually does the work is provenance. Every piece of knowledge the system holds has to carry where it came from; the source, its standing, how it relates to the others&#8230; and there has to be a rule for which source wins when they conflict. Do that, and &#8220;current&#8221; stops meaning &#8220;latest&#8221; and starts meaning &#8220;authoritative.&#8221; That&#8217;s the whole difference between a memory system and a confident guess.</p><h4>Where Provenance isn&#8217;t enough</h4><p>And where provenance runs out, where two documents genuinely conflict and no rule resolves it, the honest move is the same one the rest of the box already makes. Don&#8217;t pick the winner and hope. Surface both, cited with provenance attached, and hand the call to the person whose call it actually is. A system that says &#8220;these two conflict, and here&#8217;s where each came from&#8230; you decide&#8221; is worth more than one that confidently quotes last year&#8217;s policy because it happened to sort first.</p><p>There&#8217;s an ownership point buried in here too. A memory you can&#8217;t trace isn&#8217;t really yours&#8230; it&#8217;s a pile you&#8217;re trusting on faith. Knowing where each thing came from, and who says it&#8217;s still true, is part of what it means to own your data instead of renting a vendor&#8217;s black box that remembers whatever it remembers. It&#8217;s how humanity has always ascertained ground truth.</p><p>Recency isn&#8217;t authority. Provenance is. Everything else is a nicer label on the same guess.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://research.thirdshiftrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Local A.I, Applied. blog by ThirdShift R&amp;D! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[02 - The cheapest local AI box that actually works]]></title><description><![CDATA[The parts, the honest numbers, what fought me, and where it falls short]]></description><link>https://research.thirdshiftrd.com/p/the-cheapest-local-ai-box-that-actually</link><guid isPermaLink="false">https://research.thirdshiftrd.com/p/the-cheapest-local-ai-box-that-actually</guid><dc:creator><![CDATA[Local A.I. Applied]]></dc:creator><pubDate>Thu, 23 Jul 2026 18:38:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5b11ba22-87b5-405e-b22e-906cb098ba45_1456x819.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>[drafted by a project agent, edited by me, by hand]</p><p>I wanted a local AI box the way I&#8217;d want any tool in the shop: owned, understood, serviceable by me. No subscription, no cloud, nothing phoning home. So I built the cheapest one I could that still does real work all day. Not the fastest box&#8230; the cheapest one that&#8217;s actually viable. Here&#8217;s the whole thing, with the real numbers and the parts that fought me.</p><h3><strong>The shape of it</strong></h3><p>A Ryzen 9 mini-PC (mine&#8217;s a Minisforum UM890 Pro) which stays a normal everyday computer, and an Arc Pro B60 24GB rides alongside it on an open-air eGPU dock with its own 650W Gold PSU, connected over OCuLink. Two machines wearing one job. The mini-PC never gives up its own silicon to the model, and if the dock ever unplugs you&#8217;ve still got a first-rate desktop. If your host has no native OCuLink port, you can enumerate the card off an M.2-to-OCuLink adapter, just keep the cable short.</p><h3>The part that fought me</h3><p>Getting the card to enumerate cleanly meant sorting a small-BAR issue during bring-up. Once that was handled it comes up clean every boot. The link is Gen4-capable and I&#8217;ve watched it train there after a re-enumeration; this boot is sitting at Gen3 x4 under load, and every number in this post reproduces at Gen3. the weights live in VRAM, so the link isn&#8217;t your bottleneck either way.</p><h3>The real lesson: it&#8217;s the model, not the card</h3><p>I came from a dense 32B doing 15.5 tok/s and it was a slog. I swapped to a 35B-A3B MoE and the same card runs it fully offloaded: about 42 tok/s at low context on my daily serving config (q8 KV cache + flash attention), and around 8K depth it holds 29&#8211;32 tok/s. (A shallow-KV bench config reads higher, ~45 I&#8217;m quoting the as-served number on purpose, because a tok/s figure without its config is half a number.) At 8K the gap gets embarrassing: about 30 versus the dense 32B&#8217;s 7.15. better than 4x, because only ~3B parameters are active per token. If you&#8217;re budget-building, the MoE is the upgrade; the GPU is just the enabler.</p><p>First-hand Arc gotcha while you&#8217;re benching: a q8 KV cache makes flash-attention mandatory on SYCL, and -fa auto will silently switch it on. I briefly believed I had &#8220;flash-attn-off&#8221; numbers that turned out not to exist. Check what actually resolved.</p><h3>The surprise</h3><p>The bigger model holds more context on the same card. The 35B MoE fits 32K and still moves at ~16 tok/s, while the dense 32B flat OOMs at 32K. That&#8217;s measured, its memory footprint at depth is far smaller. I&#8217;m telling you what the meter showed, not theorizing why.</p><h3>It does three jobs at once</h3><p>The 35B MoE lives on the B60, a vision model rides the Ryzen&#8217;s iGPU for paper and photos, and a voice model runs on two CPU cores with zero VRAM. On the vision. the honest limit, it does not read values off a chart or a dense number grid. That hit an unimplemented op on my Arc/SYCL stack, so the box hands the chart back to you instead of guessing at it. Nothing shares, nothing waits.</p><h3>Thermals, measured because I don&#8217;t trust vibes</h3><p>Half an hour flat out: 60 to 62&#176;C, never throttled, on the open-air dock.</p><p>Power, honestly</p><p>Across the whole box, wall draw runs 55 to 190W idle to load, with a loose ballpark around 145W when it&#8217;s working. That&#8217;s the range, not a budget, the detailed efficiency picture. What a single answer actually costs, is its own measurement. I&#8217;d rather run the wall meter properly and show the method than ballpark it here. That writeup&#8217;s coming.</p><h3>The honest limits</h3><p>I won&#8217;t pretend Intel cards are competitive for gaming, and the software stack is younger than the green team&#8217;s, you&#8217;ll read more docs getting to the same places, and work harder through issues that are novel. But for the money, 24GB of VRAM from a $650ish GPU that runs a 30B-class brain all day is a hard argument to lose.</p><h3>Money</h3><p>If you already own an OCuLink mini-PC, the card, dock, PSU and cable land under a grand. Buying everything new including the host, figure $1,500 to $1,800 in parts at June 2026 prices. One thing I want to be straight about, that number does not include RAM or an SSD. Those ride whatever memory and storage prices are doing the day you buy, and they swing. So, budget them on top rather than getting surprised at checkout. The dedicated AI appliances in this class run $4,000+ and only do the AI half.</p><p>For a box meant to run continuously, I cap self-scheduled background work at 80% of idle&#8230; that&#8217;s a warranty-and-margin call, not a thermal one. Nothing in this build overheats at full load; 80% just keeps it unambiguously inside &#8220;normal use&#8221; and leaves headroom to catch up.</p><p>Owned, understood, serviceable, and nothing leaves the building. That was the whole point.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://research.thirdshiftrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Local A.I, Applied. blog by ThirdShift R&amp;D! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[01 - A cited answer is not a correct answer]]></title><description><![CDATA[Our benchmark sat green for weeks while the model reasoned right and wrote the conclusion wrong. The measurement, and what we changed.]]></description><link>https://research.thirdshiftrd.com/p/01-a-cited-answer-is-not-a-correct</link><guid isPermaLink="false">https://research.thirdshiftrd.com/p/01-a-cited-answer-is-not-a-correct</guid><dc:creator><![CDATA[Local A.I. Applied]]></dc:creator><pubDate>Mon, 20 Jul 2026 11:47:40 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3c8e6603-8c9d-409b-b452-3d64db1d6b78_1456x819.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We build a local AI that reads a small business&#8217;s own documents and answers questions about them on hardware that stays in the building, cited to the page, nothing leaving the site. The whole promise rests on one word: cited. If it tells you something, it shows its source, so your people can check it.</p><p>For weeks, our benchmark told us that promise was holding. One column in particular synthesis questions, where the answer has to be assembled from a couple of figures in the source sat green. Right numbers, right citations, every time. </p><p><em>The column was lying to us</em>. We just weren&#8217;t asking it the right question.</p><h4><strong>What we were actually measuring</strong></h4><p>Our scorer, on those synthesis questions, checked two things: did the answer surface the correct figures, and did it bind them to the correct citation. Both were true. So, it scored PASS.</p><p><em>What it never did was read the conclusion.</em></p><p>Here&#8217;s the failure that hid behind the green, on a filing-threshold question &#8212; the kind a real owner actually asks. The model retrieved the right threshold. It cited the right page. In its own words, it did the arithmetic correctly: &#8220;since $500 is above the $400 threshold, they are generally required to file.&#8221; And then, one line later, it wrote the answer as: &#8220;(a) No.&#8221;</p><p>It reasoned its way to the right answer and then wrote down the wrong one. Because the figure and the citation were both present and both correct, our benchmark called it a clean pass.</p><h4>The number, honestly</h4><p>We built a separate axis that reads the stated conclusion against what the cited facts actually entail and measured it. On that one inference question, across depths and positions, the model&#8217;s stated yes/no answer was correct 44% of the time (Wilson 95%: 34.6&#8211;54.7%). A large chunk of the rest weren&#8217;t random errors &#8212; they were answers that reasoned correctly in prose and then labeled the conclusion backwards.</p><p>The honest limits: this is one question type so far, one corpus, one model. It is not &#8220;the model is wrong 44% of the time.&#8221; It is &#8220;on questions that need a step of judgment, the stated conclusion decouples from the reasoning, and we&#8217;ve now measured it once.&#8221; We&#8217;re expanding it to more inference types before we&#8217;d quote a general rate. But one clean measurement of this was enough to change how we think about the whole product.</p><h4>Why this is worse than a plain error</h4><p>A model that says &#8220;I don&#8217;t know&#8221; costs you nothing. A model that hallucinates an obvious nonsense answer, you catch. But a model that lays out the correct reasoning, cites the correct page, and then states the wrong conclusion has done the most dangerous thing an assistant can do; it has spent the trust its visible work just earned. The citation is what makes you stop checking. The reasoning is what makes you believe. And then the answer is wrong.</p><p>For a tool whose entire pitch is &#8220;cited so you can verify,&#8221; that is the one unforgivable failure. It punishes exactly the person who trusted it correctly.</p><h4>The lesson, if you&#8217;re evaluating RAG</h4><p>Check whether your scorer reads the conclusion or just the retrieval. &#8220;Grounded&#8221; and &#8220;correct&#8221; are different axes, and a system can be flawless on the first while failing the second. If your eval only measures whether the right chunk was retrieved and cited, you can ship a green dashboard on top of a product that reasons right and answers wrong &#8212; and you won&#8217;t find out from your metrics. You&#8217;ll find out from a customer who trusted a cited number.</p><p>We only caught it because we started scoring the model against its own stated reasoning. The tell was cheap once we looked: reasoning says one thing, label says the opposite, in the same answer.</p><h4>What we did about it</h4><p>We didn&#8217;t try to make the model conclude correctly and call it fixed. We changed what the product is allowed to do.</p><p>On a question that requires a judgment call, for example &#8220;do I have to file?&#8221;  as opposed to &#8220;what is the filing threshold?&#8221; the box doesn&#8217;t render the verdict. It surfaces the retrieved facts and the governing rule, cited to the page, and hands the call to the person whose call it actually is&#8230; the owner or responsible employee. It files, it cites, it assembles the whole context on demand and then it says, in so many words, &#8220;this one&#8217;s a judgment call, and it&#8217;s yours.&#8221;</p><p>That&#8217;s not a limitation we&#8217;re apologizing for. It&#8217;s the design. The people we build for are good at making calls on incomplete information, that&#8217;s their job. What they don&#8217;t have time for is assembling the context. So, we do the half a machine is good at; hold everything, cite exactly, never tire&#8230; and we hand off cleanly at the seam: the exact point where a human&#8217;s judgment belongs and a model&#8217;s confidence is a liability.</p><p>The benchmark being green taught us less than the benchmark being caught. We&#8217;ll keep publishing both. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://research.thirdshiftrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Local A.I, Applied. blog by ThirdShift R&amp;D! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Start Here]]></title><description><![CDATA[One industrial tradesman building local AI you can actually see, showing the work: the tests, the honest numbers, and the failures.]]></description><link>https://research.thirdshiftrd.com/p/start-here</link><guid isPermaLink="false">https://research.thirdshiftrd.com/p/start-here</guid><dc:creator><![CDATA[Local A.I. Applied]]></dc:creator><pubDate>Mon, 20 Jul 2026 11:25:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DPkx!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f4be89e-bfa6-4492-b84f-7b7c9adce210_2688x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>ThirdShift R&amp;D is one industrial electrician building AI that runs where you can see it.</p><p>We&#8217;re not a frontier lab. We&#8217;re not training the next giant model or chasing a benchmark record ; the frontier builds the engine, and that&#8217;s a different trade than ours. Ours is the unglamorous work in the middle: getting these systems across the gap from a demo that impresses researchers to something a regular business can actually run&#8230; safely, in the hands of someone who isn&#8217;t an AI expert and shouldn&#8217;t have to be.</p><p>That gap is where almost everything fails. A model can reason like a PhD and still get wired into a workplace in a way that quietly burns the people trusting it. Closing it and making powerful, unpredictable systems safe for regular people doing regular work is a real trade. It&#8217;s the one we practice. It&#8217;s the one the human behind this page paid the bills at his day job, handling voltages with consequences to a human body that can be just as fatal as a bad A.I. may be to an entire organization.</p><p>The product is a box that sits on your site, reads your own documents, and answers cited to the page and then hands the judgment calls back to you, because those are yours to make. Nothing leaves the building. No call ever made without a Human-in-the-Loop.</p><p>This is where we show the work. Not a sales pitch, not a benchmark nobody can reproduce&#8230; the actual tests, the methods, the numbers with honest error bounds, and the failures, especially the failures. When we catch our own tools lying to us, that goes here too, with the number attached.</p><p>If you build or deploy local AI or you&#8217;re an operator trying to decide whether any of this is trustworthy enough to put near your business&#8230; Well, this is for you. If you&#8217;re here for &#8220;AGI next quarter,&#8221; you&#8217;ll be bored. If you want passed the hype, and to start looking at how this technology is being applied on-site, and how it is likely to affect YOUR everyday workplace, you are in the right place. </p><p>We leave no doubt. Including about where we were wrong. Follow along.</p><p>&#8212; ThirdShift R&amp;D</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://research.thirdshiftrd.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Local A.I, Applied. blog by ThirdShift R&amp;D! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>