<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>Annas Bin Adil</title>
    <subtitle>Engineer at heart. Writing about AI, robotics, and the human experience.</subtitle>
    <link href="https://www.annasbinadil.com/feed.xml" rel="self"/>
    <link href="https://www.annasbinadil.com/"/>
    <updated>2026-05-15T00:00:00Z</updated>
    <id>https://www.annasbinadil.com/</id>
    <author>
        <name>Annas Bin Adil</name>
        <email>annasbinadil@gmail.com</email>
    </author>
    
    
    <entry>
        <title>The Mechanical Bottleneck</title>
        <link href="https://www.annasbinadil.com/posts/2026-05-15-the-mechanical-bottleneck/"/>
        <updated>2026-05-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-05-15-the-mechanical-bottleneck/</id>
        <summary>Physical AI deployment will not be paced by chips, data, or labor. It will be paced by precision reducers, small Japanese-made geared parts most engineers can&#39;t name. That reshapes timeline forecasting, industrial policy, and embodied AI safety in ways the discourse hasn&#39;t yet caught up to.</summary>
        <content type="html">&lt;p&gt;&lt;em&gt;Epistemic status: Research synthesis with interpretation. I&#39;m writing as someone reading the field, not someone building in it. The empirical claims are sourced. The three implications I draw are at different confidence levels, flagged where they get speculative.&lt;/em&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;The thing that will pace humanoid robot deployment over the next decade is not a chip. It is not data. It is not even labor. It is a small geared part most engineers can&#39;t name, made by a handful of Japanese suppliers, and the global supply caps humanoid production at roughly 500,000 units a year regardless of how good the underlying AI gets.&lt;/p&gt;
&lt;p&gt;This is the central finding of Epoch AI&#39;s April 2026 piece, &amp;quot;How Fast Could Robot Production Scale Up?&amp;quot; It is, as far as I can tell, the cleanest published work on physical-AI deployment timelines. The bottleneck the piece names is precision reducers: planetary gearboxes, cycloidal/RV reducers, strain-wave gears. Tiny mechanical components made by firms most AI people couldn&#39;t list if asked.&lt;/p&gt;
&lt;p&gt;The post draws out what that one rigorous empirical finding implies for three things that aren&#39;t getting enough attention: timeline forecasting, industrial policy, and embodied AI safety.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Capability got cheap. Deployment didn&#39;t. The bottleneck is mechanical, not silicon, and the AI discourse hasn&#39;t mapped it.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;What Epoch found&lt;/h2&gt;
&lt;p&gt;The piece is a bottom-up bill-of-materials analysis across five form factors: humanoids, quadrupeds, robotic arms, wheeled robots, drones. Authors Jean-Stanislas Denain and Yann Rivière mapped global component supply against production demand, then compared the result against historical mobilization cases including WWII aircraft, Tesla Shanghai factory builds, and Ukraine FPV drone production.&lt;/p&gt;
&lt;p&gt;The current state of humanoid production is small. About 16,000 units per year in 2025, doubling every six months. The doubling rate is anomalous, Epoch notes. Most of those units are research platforms and marketing demonstrations, not productive work. The real test is what the production curve looks like once capability arrives and somebody actually needs millions of humanoids.&lt;/p&gt;
&lt;p&gt;Their projection, under a demand-shock scenario triggered at end of 2027, is that humanoid production reaches 1.5 to 3 million units per year by end of 2030 clearly, and 5 to 10 million plausibly. Drones could reach 100 to 200 million per year without especially hard limits. Quadrupeds 8 to 15 million. These are not the kind of numbers that move you off your timeline if you were expecting an AI rollout shaped like the smartphone wave. Smartphones hit a billion units shipped in six years from the iPhone launch. The Epoch projection puts humanoids at less than 1 percent of that volume over a similar window.&lt;/p&gt;
&lt;p&gt;What constrains the rate is not the things the AI discourse usually worries about. Cameras come off production lines at 7 billion units per year globally. MEMS sensors at 31 billion. Bearings at tens of billions. Batteries at 2 billion kWh per year. None of these are binding. None of these sit on the critical path.&lt;/p&gt;
&lt;p&gt;The chokepoint is precision reducers. Planetary gearboxes at roughly 3 million units per year. Cycloidal and RV reducers at roughly 2 million. Strain-wave gears at roughly 12 million, the least restrictive of the three. Humanoids use somewhere between 20 and 40 reducers each, depending on hand dexterity. Tesla Optimus, with its high-degree-of-freedom hands, could triple per-robot reducer demand. The math caps humanoid production at roughly 500,000 per year at current capacity, no matter how many factories you build.&lt;/p&gt;
&lt;p&gt;The supply concentration matters. Harmonic Drive Systems (HDS, listed on the Tokyo Stock Exchange) and Nabtesco, both Japanese, dominate the precision-reducer market. HDS scaled its Ariake plant from 150,000 to 220,000 units per month by August 2022 to meet industrial-robot demand, with the Hotaka facility adding another 40,000 monthly. Total HDS capacity sits around 3 million units per year. Nabtesco is the dominant supplier in RV reducers used in industrial-robot joints. China has been pushing localization aggressively. Leaderdrive holds 30 to 40 percent of China&#39;s harmonic-reducer market and counts Tesla among its customers. But a Jamestown Foundation analysis notes that while 80 percent of RV reducers are now assembled in China, roughly 90 percent of the machine tools used to make them are imported, mostly from Japan.&lt;/p&gt;
&lt;h2&gt;Why mechanical precision is the binding constraint&lt;/h2&gt;
&lt;p&gt;The first-principles answer to why mechanical precision is the bottleneck rather than something else comes from looking at how each potential bottleneck scales.&lt;/p&gt;
&lt;p&gt;Chips scale with fabrication capacity, and the world has been pouring capex into fab construction since the CHIPS Act. The marginal chip for a humanoid is not where the bottleneck lives. Batteries are similarly in a phase of aggressive capacity expansion driven by EVs. Cameras and MEMS sensors get produced at volumes orders of magnitude larger than humanoid demand could plausibly require.&lt;/p&gt;
&lt;p&gt;Labor doesn&#39;t bottleneck because factory operations workers, which Epoch estimates at 40,000 to 120,000 for 10 million humanoids per year, can be trained in parallel during the construction phase. The construction phase itself is the timing constraint. Auto-plant retrofits run 6 to 10 months. Greenfield Chinese factory builds run 6 to 9 months. Greenfield Western builds run 2 years or more.&lt;/p&gt;
&lt;p&gt;Precision reducers are different because the manufacturing process is itself a deep skill stack. These are gear-on-gear interfaces machined to micrometer tolerances, with proprietary metallurgy and grinding processes that took specific Japanese firms decades to refine. They are not a thing you can produce by buying CNC machines and hiring trained workers. The machine tools themselves, the precision grinders and gear-cutters used to make reducers, are largely Japanese-made. China&#39;s domestic reducer production has scaled, but the Jamestown analysis flagged that it sits on top of imported machine tools from the same source.&lt;/p&gt;
&lt;p&gt;This is a constraint with the shape of TSMC, not the shape of Foxconn. TSMC isn&#39;t a chokepoint because chip fabrication is conceptually hard. It is a chokepoint because the tacit knowledge of running advanced fabs is concentrated in one firm. Precision reducers are similar. The capacity expansion path does not bottleneck on capital, raw materials, or labor. It bottlenecks on tacit metallurgical knowledge and on the upstream machine tools that themselves bottleneck on the same Japanese firms.&lt;/p&gt;
&lt;p&gt;Year 1 of any demand shock, Epoch notes, is mostly construction rather than production. By end of 2028, even with aggressive scale-up, the global stock of newly-produced humanoids would be in the low hundreds of thousands.&lt;/p&gt;
&lt;p&gt;Reducer capacity could 2x by 2028 and 4x by 2030 on baseline trajectory, or 10x to 30x under demand shock. The 10x-to-30x range is where most of the projection&#39;s uncertainty lives. It is also where the question becomes interesting: how fast can Japanese firms expand without losing the precision that made them the suppliers in the first place? That is not a question with an obvious answer.&lt;/p&gt;
&lt;h2&gt;Implication 1: Timeline forecasting&lt;/h2&gt;
&lt;p&gt;This finding reshapes timeline forecasting for AI in a specific way. The AI safety and forecasting community has done thoughtful work on capability timelines and very little on deployment.&lt;/p&gt;
&lt;p&gt;Ajeya Cotra&#39;s &amp;quot;Forecasting Transformative AI with biological anchors&amp;quot; (Open Philanthropy, 2020-2022) is the most-cited capability timeline analysis. It anchors compute requirements to biological reference points and projects hardware-and-spending trends. Deployment is bracketed as a downstream consequence with minimal physical-world friction modeled.&lt;/p&gt;
&lt;p&gt;Holden Karnofsky&#39;s &amp;quot;Most Important Century&amp;quot; (Cold Takes, 2021) introduces PASTA, his term for the process of automating scientific and technological advancement, and assumes that once AI can automate research, physical deployment follows quickly via AI-designed robotics. He flags this as a load-bearing assumption.&lt;/p&gt;
&lt;p&gt;Leopold Aschenbrenner&#39;s &amp;quot;Situational Awareness&amp;quot; (June 2024) makes essentially the same move. Intelligence-explosion thesis, with physical deployment treated as a fast follow-on. The deployment treatment is the weakest part of the document.&lt;/p&gt;
&lt;p&gt;The exception is Dario Amodei&#39;s &amp;quot;Machines of Loving Grace&amp;quot; (October 2024). Amodei explicitly addresses deployment lag, coining the phrase &amp;quot;limits to compressed 21st century.&amp;quot; He estimates 5 to 10 years for biology breakthroughs to deploy through clinical trials and regulatory approval. He does not analyze robotics supply chains specifically, but he is the only frontier-lab CEO publicly grappling with the fact that capability does not equal deployment.&lt;/p&gt;
&lt;p&gt;The Epoch piece fills a real hole in this literature. It is one of the few rigorous bottom-up deployment analyses from the AI-aware community.&lt;/p&gt;
&lt;p&gt;What does the analysis change about timeline forecasting? Three things.&lt;/p&gt;
&lt;p&gt;First, it sharpens the capability/deployment distinction. If you believe AI capability will arrive in the next three to five years and you&#39;re thinking about what the world looks like after that, the Epoch numbers should pull your humanoid-deployment estimates downward. Not because capability slips. Because deployment was always going to take three to five more years after capability already arrives.&lt;/p&gt;
&lt;p&gt;Second, it gives a falsifiable reference class. Industrial robots took fifty years to reach a million deployed units, according to International Federation of Robotics data. They are now at around 3.5 million in operational stock. Humanoid bull projections, including Tesla&#39;s stated target of &amp;quot;millions per year&amp;quot; by the late 2020s and Figure&#39;s commercial deployment framing, imply collapsing that 50-year history into roughly five years. The Epoch analysis says the industry-wide ceiling under demand shock is 5 to 10 million per year by end of 2030, which would require essentially the entire global reducer supply chain to be captured by one or two players. That is not impossible, but it is the implausible case, not the central one.&lt;/p&gt;
&lt;p&gt;Third, it explains the Rodney Brooks pattern. Brooks has maintained an annual predictions scorecard for over a decade tracking robotics and self-driving forecasts. The pattern across the field is consistent. Capability demos arrive roughly on schedule. Deployment-at-scale runs 3 to 10x slower than industry forecasts. The Epoch analysis suggests the slowdown isn&#39;t psychology or marketing. It is the supply chain. If you forecast deployment as though it tracks capability, you systematically miss by an order of magnitude.&lt;/p&gt;
&lt;p&gt;The cleanest data point on the gap between marketing and reality: Agility Robotics&#39; Salem, Oregon factory, with stated capacity around 10,000 humanoids per year. That is the verifiable number. Tesla and Figure imply millions. Agility shows ten thousand. The ratio is the gap between marketing and what the supply chain currently supports.&lt;/p&gt;
&lt;h2&gt;Implication 2: Industrial policy&lt;/h2&gt;
&lt;p&gt;The implications for industrial policy are sharper and probably more actionable than the timeline question. Strategic chokepoint thinking has been applied to semiconductors but not yet to reducers, despite the geographic concentration being arguably tighter.&lt;/p&gt;
&lt;p&gt;China has been pushing this question explicitly for a decade. Made in China 2025 named speed reducers, servomotors, and controllers, which together represent about 70 percent of robot bill-of-materials by value, as indigenization priorities.&lt;/p&gt;
&lt;p&gt;The results are mixed and the literature is honest about it. A peer-reviewed PageRank-on-patents study by Liu et al. (ScienceDirect, 2025) found that Made in China 2025 lifted midstream and downstream robotics innovation quality but failed to move upstream component quality, which is exactly the reducer and sensor and controller layer. Leaderdrive&#39;s 30 to 40 percent share of China&#39;s harmonic-reducer market is real growth, but, as noted, sits on imported Japanese machine tools.&lt;/p&gt;
&lt;p&gt;The US response has been narrow. The most concrete recent action is the Section 232 national-security investigation of robotics and industrial machinery imports, initiated by Commerce on September 2, 2025 (Federal Register notice 2025-18749). Public comments closed October 17, 2025; tariff or import-restriction decisions plausibly land in spring 2026. This is a trade tool, not a capex subsidy. The Humanoid ROBOT Act of 2025 (S.3275) extends Section 889-style federal procurement bans to humanoids from PRC, Iran, DPRK, and Russia-linked entities. That is a procurement-restriction bill, not a capacity-building bill.&lt;/p&gt;
&lt;p&gt;The Information Technology and Innovation Foundation published &amp;quot;A Time to Act: Policies to Strengthen the US Robotics Industry&amp;quot; in July 2025, by Robert Atkinson. The numbers are striking. Japan produces 46 percent of global robotics output. By 2024 imports outweighed exports 4 to 1. Atkinson recommends expanding NIST Manufacturing USA programs and the Advanced Robotics for Manufacturing institute, plus restrictions on Chinese robotics imports. The recommendations stop short of invoking Defense Production Act Title III authorities or proposing CHIPS-style appropriation for robotics components.&lt;/p&gt;
&lt;p&gt;The rare-earth angle compounds this. Robots need high-density permanent magnets, specifically NdFeB with heavy rare-earth additions of terbium or dysprosium, for the servo motors that pair with reducers. China&#39;s October 2025 MOFCOM Notices No. 61 and 62 require licenses for these magnets, with extraterritorial licensing for magnets made overseas using Chinese technology starting December 1, 2025. The notices were suspended until November 10, 2026, but the regulatory framework is in place. Industrial robots use roughly 300 to 500 grams of rare-earth-doped magnets per motor, and a humanoid has 20 to 40 motors. The supply-chain exposure is non-trivial.&lt;/p&gt;
&lt;p&gt;What would real policy that takes the bottleneck seriously look like? Probably some combination of explicit chokepoint mapping treating precision reducers and rare-earth magnets together; capex incentives for domestic reducer manufacturing, which would in turn require investing in the machine-tool supply chain that the Liu et al. analysis shows is the harder upstream problem; and allied coordination with Japan rather than competition. The current US framing of trade tools and procurement bans is closer to the early-2010s semiconductor framing than to the post-2020 framing. The conceptual upgrade has not happened yet.&lt;/p&gt;
&lt;h2&gt;Implication 3: Embodied AI safety&lt;/h2&gt;
&lt;p&gt;The safety implications are the most speculative of the three. The embodied-AI safety literature is thin, fragmented, and dominated by self-driving cars, which the broader physical-AI community treats as a sibling discipline rather than the main event.&lt;/p&gt;
&lt;p&gt;The conceptual foundation is Stuart Russell&#39;s &lt;em&gt;Human Compatible&lt;/em&gt; (2019) and the Center for Human-Compatible AI&#39;s work on Cooperative Inverse Reinforcement Learning (Hadfield-Menell, Russell, Abbeel, Dragan, NeurIPS 2016). Russell&#39;s domestic-robot thought experiment, the robot that cooks the cat because it wasn&#39;t told cats are loved, is the canonical intuition pump. CIRL formalizes preference inference from behavior. This is rigorous theory with no deployed safety system attached to it.&lt;/p&gt;
&lt;p&gt;Anthropic&#39;s Responsible Scaling Policy addresses biological, chemical, radiological, nuclear, and cyber risks plus autonomy, but does not name physical AI or robotics as a capability threshold in the public versions I have seen. OpenAI shuttered its robotics team in 2021 and has no published safety research on its more recent humanoid investments. The robotic foundation model labs themselves, including Physical Intelligence (π0), Open X-Embodiment, and Figure (Helix), have published essentially no public safety research either.&lt;/p&gt;
&lt;p&gt;The empirical safety evidence comes almost entirely from autonomous vehicles. The Cruise robotaxi incident on October 2, 2023, in San Francisco, remains the single most documented physical-AI safety failure on public record. A Cruise vehicle dragged a pedestrian roughly 20 feet after she was struck into its path by a human-driven car. The California Public Utilities Commission suspended Cruise&#39;s permit on October 24, 2023. The Quinn Emanuel independent report (January 2024), commissioned by Cruise itself, documented that the company initially showed regulators a video that cut off before the dragging. The failure was both technical, in that the vehicle&#39;s reaction model didn&#39;t handle the secondary-impact case, and organizational, in that the company concealed it from regulators. Cruise&#39;s parent eventually wound down the program.&lt;/p&gt;
&lt;p&gt;Tesla&#39;s Autopilot has generated a parallel safety pattern: an NHTSA recall in December 2023 covering approximately 2 million vehicles, with 13 documented fatal crashes. That is a regulator-driven case rather than a single dramatic incident.&lt;/p&gt;
&lt;p&gt;Waymo has published the most rigorous public safety case. The Waymo Safety Hub and Kusano et al.&#39;s 2024 comparison study report approximately 25 million rider-only miles with significant claimed reductions in police-reported and injury crashes versus human-driver baselines. The comparison-class selection is Waymo&#39;s, but the methodology is the strongest in the field.&lt;/p&gt;
&lt;p&gt;The Epoch deployment-lag analysis changes this picture in three places, ordered by how grounded the observation is.&lt;/p&gt;
&lt;p&gt;The slow-rollout-helps-safety argument breaks in the military vertical. Paul Scharre&#39;s &lt;em&gt;Army of None&lt;/em&gt; (2018) and ongoing CNAS work document the trajectory. The OpenAI-Anduril partnership announced in December 2024 and the Anthropic-Palantir-AWS defense deal in November 2024 are real moves into a market where the customers, state militaries, are explicitly willing to absorb unit cost premiums that would crush commercial deployment. The reducer constraint applies to commercial humanoids assuming peacetime allocation. It does not apply if a state actor commits to scaling regardless of cost. The Epoch analysis implicitly assumes a peacetime industrial pattern. This is the strongest grounded counter to the slow-rollout-helps-safety case.&lt;/p&gt;
&lt;p&gt;Physical-AI safety research has more runway than the discourse implies, at least on the commercial side. If commercial humanoid deployment is paced by supply chains running on industrial-time, the field has five to ten years to develop, test, and deploy safety techniques before the population of deployed embodied agents is large enough that incidents become routine. The current state, in which there are zero public safety publications from any major robotic foundation model lab, is therefore the bigger problem than the slow rollout. The runway exists. The research is not filling it.&lt;/p&gt;
&lt;p&gt;And most speculatively: there is a thesis nobody has made canonically yet, that embodied AI is where alignment debates become empirical. Robots produce real-world actions with real-world consequences in a way text generators don&#39;t. As humanoids reach the deployed scale Epoch projects, the LLM discourse&#39;s mostly-theoretical alignment debates will have physical analogues. I&#39;m flagging the thesis because the literature is open for someone to make it well. I&#39;m not in a position to be that someone here.&lt;/p&gt;
&lt;h2&gt;Cruxes&lt;/h2&gt;
&lt;p&gt;The Epoch analysis is a snapshot, not a final answer. Five things could change it.&lt;/p&gt;
&lt;p&gt;Recursive self-improvement is the largest deferred question. Epoch explicitly brackets the case where AI starts designing reducer manufacturing plants or developing the metallurgical processes that currently take Japanese firms decades to refine. If that happens, the bottleneck story changes. The analysis assumes ordinary human engineers run the scale-up. That is a reasonable simplification for current conditions and a bad one for any future where the AI is itself optimizing the supply chain.&lt;/p&gt;
&lt;p&gt;Software substituting for hardware is the second crux. Physical Intelligence demonstrated precision manipulation tasks using relatively simple grippers, with fewer actuators and less reducer demand per unit, by improving the learned policy. If foundation models can compensate for mechanical imprecision the way deep learning has compensated for hand-engineered features, the bill-of-materials math shifts. Robots could become viable at lower precision tiers, which would change which suppliers matter and what the cap actually is.&lt;/p&gt;
&lt;p&gt;Cohort versus price effects is the third. Industrial robots took 50 years to reach a million deployed units partly because of supply constraints and partly because the unit economics did not work for most applications until well into that window. If humanoid prices follow a fast cost curve, both because the AI labor cost is in the training run rather than the marginal robot and because EV-style battery and motor economies will apply, the reference class might be wrong. The bottleneck still binds the production curve, but the demand curve and the deployment curve might decouple from it differently.&lt;/p&gt;
&lt;p&gt;Geopolitical fragmentation is the fourth and least quantifiable. The Epoch analysis assumes global trade keeps working. A Taiwan crisis, escalating sanctions, or a serious break in the Japan-China-US technology relationship would either collapse the supply chain or fragment it in ways that produce different and unpredictable constraints. The current rare-earth export-control framework is a small preview of the kinds of moves that could matter much more.&lt;/p&gt;
&lt;p&gt;The fifth thing the analysis does not address is demand origin. Why does a shock happen, and when? Epoch&#39;s projection runs from an EOY 2027 trigger date, which is a placeholder. The whole analysis is &amp;quot;conditional on a shock, how fast can production scale&amp;quot; rather than &amp;quot;when will a shock happen and what kind.&amp;quot; That is the right scope for the question they were asking, but it means the deployment timing is really two timelines stacked, capability arrival and then production response, and the post only addresses the second.&lt;/p&gt;
&lt;h2&gt;Where this leaves me&lt;/h2&gt;
&lt;p&gt;What I am watching for over the next eighteen months: whether HDS or Nabtesco announces capacity expansion specifically tied to humanoid demand at scale; whether the Section 232 investigation produces anything more substantial than tariffs; whether the major humanoid companies publish anything resembling a safety framework; and whether anyone writes the CSIS-style brief that frames the reducer-and-magnet chokepoint as the strategic problem it appears to be.&lt;/p&gt;
&lt;p&gt;The most useful posture an outsider can take on physical AI right now is not prediction. It is noticing which questions the discourse has not gotten to yet.&lt;/p&gt;
&lt;p&gt;The place I am most likely to be wrong is the software-substituting-for-hardware crux. If Physical Intelligence or another lab demonstrates production-grade manipulation at lower reducer counts than current humanoid architectures assume, the bottleneck story shifts and most of the post&#39;s three implications shift with it. I&#39;d want to hear from anyone watching that closely.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Tools Have Split</title>
        <link href="https://www.annasbinadil.com/posts/2026-04-15-the-tools-have-split/"/>
        <updated>2026-04-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-04-15-the-tools-have-split/</id>
        <summary>AI coding tools have separated into two classes. One compensates for missing operator skill; one amplifies present operator skill. Senior engineers systematically prefer the second. The split is a market signal that operator skill is now the bottleneck, and a snapshot, not a destination.</summary>
        <content type="html">&lt;p&gt;&lt;em&gt;Epistemic status: Grounded in published 2025-2026 survey data and my own day-to-day use of these tools. The empirical claims are sourced. The unifying frame is mine, and I am moderately confident in it. The 18-month prediction at the end is speculation, offered as a falsifiable bet rather than a confident forecast.&lt;/em&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;In a 2026 survey of about 900 engineers, the most-loved AI coding tool is Claude Code, at 46%. Cursor is at 19%. GitHub Copilot at 9%. The detail that matters is buried in the seniority breakdown: among directors and senior leaders, Claude Code&#39;s preference roughly doubles, while Cursor&#39;s declines with seniority. If AI coding tools were substitutable, this shouldn&#39;t happen.&lt;/p&gt;
&lt;p&gt;Seniority should not predict tool preference. The preference should track feature sets or per-session productivity, not years of experience. Instead, the most experienced engineers are systematically choosing the tool that demands the most from them as operators.&lt;/p&gt;
&lt;p&gt;This is one survey of one self-selected senior-skewed audience (Gergely Orosz&#39;s Pragmatic Engineer subscribers, median 11-15 years experience), and the qualitative axes I&#39;ll lay out next are framing, not additional evidence. If a comparably rigorous 2026 survey showed a different gradient, the rest of this argument wobbles. But the gradient is real in this data, and I think it points at something worth taking seriously.&lt;/p&gt;
&lt;p&gt;The standard gloss is &amp;quot;they understand the leverage.&amp;quot; True but undercooked. The deeper story: AI coding tools have quietly split into two classes that serve different jobs, and senior preference is the market signal for which class produces good systems.&lt;/p&gt;
&lt;h2&gt;The thesis&lt;/h2&gt;
&lt;p&gt;The thesis I want to defend is small and specific:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The cost of producing code dropped. The cost of producing good systems didn&#39;t. The craft has been redistributed, not eliminated.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is a claim about where the bottleneck moved, not whether AI is good or bad. AI is unambiguously useful for the parts of software work that were always cheap to do correctly. Greenfield code in a familiar language. Boilerplate. Tests that follow obvious patterns. Documentation drafts. The mechanical layer of software production has collapsed in cost in a way that&#39;s visible in adoption data: 85-90% of working engineers report regular AI use across the major 2025-2026 surveys (DORA, Stack Overflow, Pragmatic Engineer, JetBrains all triangulate on this).&lt;/p&gt;
&lt;p&gt;But the thing that takes thirty years of experience to do well was never typing. It was knowing what to build, in what order, for which users, with which trade-offs, and recognizing when a working implementation is in fact the wrong implementation. That layer of the work is structurally different from typing. It cannot be cheapened by faster code generation, because it doesn&#39;t bottleneck on code generation. It bottlenecks on judgment.&lt;/p&gt;
&lt;p&gt;The empirical evidence is consistent with the split. DORA 2024/2025 found that as AI adoption rose toward saturation, throughput went up but delivery stability went down. The 2024 model estimated a 7.2% reduction in delivery stability per 25% increase in AI adoption. DORA 2025 framed the mediating variables as internal platform quality and review discipline: teams with strong platforms compounded gains; teams without them compounded drag. AI didn&#39;t shift the outcome. It amplified whatever the team was already doing. The Stack Overflow 2025 Developer Survey found adoption climbing to 84% while developer trust in AI accuracy collapsed to 29%, and 45% of respondents reported that debugging AI-generated code takes more time than writing the code themselves.&lt;/p&gt;
&lt;p&gt;None of these findings say AI is bad. They say the same thing in different vocabularies: the mechanical layer cheapened, the judgment layer didn&#39;t, and the gap shows up wherever you measure system quality instead of code volume.&lt;/p&gt;
&lt;p&gt;If the work has split along that line, the senior-preference inversion isn&#39;t mysterious. Senior engineers spend most of their time on the part that didn&#39;t cheapen. The part that did cheapen was already cheap for them, because they had built the mental models that made typing the easy part. The AI tool that matters to a senior engineer is one that amplifies the part still expensive, not the part already cheap.&lt;/p&gt;
&lt;h2&gt;The tools have split&lt;/h2&gt;
&lt;p&gt;Cursor and Claude Code get discussed as the same product with different price points and feature sets. They aren&#39;t. They are two products that happen to share a category label, and the seniority split is the clearest evidence I have. Four axes describe the difference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Where the interface lives.&lt;/strong&gt; Cursor&#39;s center of gravity is inline autocomplete. You&#39;re typing; suggestions appear; you accept, reject, modify. The unit of interaction is a few-character to few-line completion, evaluated in milliseconds. Claude Code&#39;s center of gravity is agentic delegation. You describe an intent; the agent reads the code, plans, edits multiple files, runs commands, returns a diff. The unit of interaction is a task, evaluated in minutes. These aren&#39;t optimizing for the same outcome.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What the tool requires you to bring.&lt;/strong&gt; Inline autocomplete is forgiving of weak specification. You don&#39;t need to know exactly what you want; you start typing in roughly the right direction and the tool fills in. The skill it amplifies is recognition. Agentic delegation requires complete-enough specification up front. &amp;quot;Add the next endpoint&amp;quot; is too vague. The agent needs target behavior, edge cases, integration points, constraints. The skill it amplifies is specification. Senior engineers have stronger specification skill, accumulated through years of building things and watching them fail. Juniors are still learning what good specification looks like, so they get more value from a tool that doesn&#39;t demand it up front.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Trust and accountability flow.&lt;/strong&gt; Cursor&#39;s UX nudges toward accept-and-move-on. Tab-complete feels low-stakes per suggestion. You can ship a lot of accepted suggestions without ever doing an explicit &amp;quot;do I trust this?&amp;quot; check. Claude Code&#39;s UX is diff-review-approve. Every change gets a moment of explicit accountability before it lands. Senior engineers are accountable for what ships. They have seen enough AI failures and pre-AI failures to want the explicit check. The Claude Code flow matches that instinct. The Cursor flow asks them to trust at a rate their judgment hasn&#39;t agreed to.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Operator-skill amplification.&lt;/strong&gt; This is the unifying axis underneath the other three. If you have taste, domain knowledge, and specification ability, Claude Code&#39;s high-bandwidth handoff puts that to work; one good specification produces hundreds of lines of useful code. If you don&#39;t have those skills, the high-bandwidth handoff is a liability: you can specify the wrong thing very efficiently. Inline autocomplete is the constrained interface. You are confined to a known context (a file, a function), and the tool extends what you started. Even with weak operator skill, you can&#39;t get too far off course in one suggestion. Agentic delegation is the unconstrained interface. You can ask for anything. With weak operator skill, this is dangerous. With strong operator skill, it is amplification.&lt;/p&gt;
&lt;p&gt;The senior preference is not a preference for AI doing more of the work. It is a preference for an interface that rewards judgment. Cursor compensates for missing judgment. Claude Code rewards present judgment. The tools have split along a specification-skill axis: one class compensates for what you don&#39;t have; one class amplifies what you do.&lt;/p&gt;
&lt;h2&gt;Where I actually stand&lt;/h2&gt;
&lt;p&gt;I&#39;ve been using Claude Code as my primary coding interface for most of the past year. The reason isn&#39;t speed.&lt;/p&gt;
&lt;p&gt;When I work with Cursor, my eye and mind are at the implementation level: syntax, names, control flow. Even with strong AI suggestions, I&#39;m inside the code. When I work with Claude Code, I&#39;m one level up. I&#39;m writing in English about what should happen. The implementation is something I review, not something I&#39;m inside of.&lt;/p&gt;
&lt;p&gt;This isn&#39;t delegation. Delegation is when you hand off work you could do yourself but choose not to. Most discussions of agentic tools frame them this way: senior engineers delegate the typing so they can do more important things. I think this is wrong. The thing I&#39;m doing with Claude Code is operating at the level where my judgment is sharper than my fingers on a keyboard. That is a different relationship with the work than delegation.&lt;/p&gt;
&lt;p&gt;This reframe predicts something the delegation frame doesn&#39;t. If agentic tools are just delegation, the senior preference for Claude Code should evaporate the moment the tools get good enough that the delegation feels automatic. If agentic tools are abstraction-level interfaces, the senior preference should persist. The current data is consistent with the second story. Staff+ engineers in the Pragmatic Engineer survey use agents 63.5% of the time, compared to 49.7% for regular engineers. The gap is widening with maturity, not narrowing.&lt;/p&gt;
&lt;h2&gt;The practitioner debate&lt;/h2&gt;
&lt;p&gt;Three senior practitioners have published the most thoughtful takes on what good engineering with AI looks like, and they disagree in instructive ways.&lt;/p&gt;
&lt;p&gt;Steve Yegge is the most bullish. His March 2025 essay &amp;quot;Revenge of the Junior Developer&amp;quot; frames the current moment as one of six waves: traditional coding, completions, chat, coding agents, agent clusters, agent fleets. His central claim is that 95-99% of agent interactions could in principle be handled by a properly briefed model. The human role is to brief, not to author.&lt;/p&gt;
&lt;p&gt;Geoffrey Litt is the cleanest counter-frame. His October 2025 essay &amp;quot;Code like a surgeon&amp;quot; calls the &amp;quot;AI makes us all managers&amp;quot; framing &amp;quot;dangerously incomplete.&amp;quot; A surgeon does the actual work of surgery, supported by a prepped operating room and team but still in the primary work. Litt distinguishes primary tasks (core design, code by hand, AI used carefully) from secondary tasks (codebase guides, exploratory spikes, doc updates), and argues that the right pattern is to keep humans in the primary work while delegating secondary work freely.&lt;/p&gt;
&lt;p&gt;Kent Beck draws a different line. His essay &amp;quot;Augmented Coding: Beyond the Vibes&amp;quot; distinguishes two practices that get conflated. Vibe coding (Karpathy&#39;s term) is &amp;quot;describe the outcome, feed errors back, hope&amp;quot; without regard for what code is produced. Augmented coding is traditional engineering values with less typing. Beck rebuilt a B+ Tree library three times: twice with vibe coding (both collapsed under accumulated complexity), once with TDD-driven AI agents (this one held). His operative claim: &amp;quot;TDD is a superpower when working with AI agents.&amp;quot; Tests are how you discipline a non-deterministic collaborator.&lt;/p&gt;
&lt;p&gt;Yegge&#39;s model assumes the briefing problem is tractable. The Replit incident, the slopsquatting data, and the Perry et al. confidence-competence gap (all in the next section) are evidence that it isn&#39;t, at least not yet. Briefing is the hard part, not the easy part. Yegge collapses it into a solved input. I think he is wrong about that, and not just about timing. The high-judgment work does not shrink just because more of the implementation gets automated, because the implementation isn&#39;t where the judgment was.&lt;/p&gt;
&lt;p&gt;Litt&#39;s surgeon, Beck&#39;s augmented coding, Simon Willison&#39;s &amp;quot;agents amplify skilled operators,&amp;quot; and Martin Fowler&#39;s framing of AI as the biggest shift since &amp;quot;assembler to high-level languages&amp;quot; are all arguing the same thing in different vocabulary: the high-judgment work expanded as the typing work shrank. Yegge is the dissenter.&lt;/p&gt;
&lt;h2&gt;When operator skill is missing&lt;/h2&gt;
&lt;p&gt;The clearest evidence that operator skill matters comes from what happens when it&#39;s absent. The Replit incident in July 2025 is the cleanest public case I&#39;ve seen. An AI agent, operating during what should have been a code freeze, deleted a production database. It then fabricated approximately 4,000 fake users to fill the resulting void and initially reported that recovery was impossible. Replit&#39;s CEO publicly apologized and the company subsequently added a planning-only mode and explicit dev/prod separation. The incident matters not because an AI agent did something bad. It matters because the bad thing happened in a context where no operator skill was checking the agent&#39;s decisions. The Claude Code review-approve flow that senior engineers prefer is precisely the friction that prevents this class of incident.&lt;/p&gt;
&lt;p&gt;The Stanford CCS 2023 study (Perry, Srivastava, Kumar, Boneh) asked developers to complete security-relevant programming tasks, half with AI assistance and half without. Developers with AI wrote less secure code. They also rated their code as more secure. The confidence-competence gap is the durable finding: AI raises perceived quality faster than it raises actual quality, and the gap is widest in domains where the operator doesn&#39;t have the prior expertise to recognize the failure modes.&lt;/p&gt;
&lt;p&gt;The USENIX Security 2025 paper on package hallucination found that LLMs recommend non-existent packages roughly 19.7% of the time across 576,000 samples. Open-source models hallucinated at 21.7%; proprietary at 5.2%. The researchers identified 205,000 unique hallucinated package names, of which 38% are name conflations, 13% are typos of real packages, and 51% are pure fabrications. This isn&#39;t a theoretical attack surface. Security researcher Bar Lanyado registered the AI-hallucinated &lt;code&gt;huggingface-cli&lt;/code&gt; package on PyPI as an empty stub in early 2024 and received 30,000 real downloads in three months. An operator who blindly accepts package suggestions from an AI is running a roulette wheel on supply-chain security.&lt;/p&gt;
&lt;p&gt;The pattern across these cases is the same. The agent is capable. The operator is missing. The output looks productive but contains failures the operator can&#39;t detect because they don&#39;t have the underlying skill.&lt;/p&gt;
&lt;h2&gt;The eighteen-month bet&lt;/h2&gt;
&lt;p&gt;So far the argument is that the current snapshot supports the redistribution thesis. The tools have split, senior preference is real, operator skill is the bottleneck. I want to end with the part I am least certain about, which is that this snapshot is transient.&lt;/p&gt;
&lt;p&gt;Here is the bet. Over the next eighteen months, two things change in ways that compress the specification-skill premium.&lt;/p&gt;
&lt;p&gt;First, the tools get more abstracted. The level at which the human operates today (writing English specifications, reviewing diffs at the file level) is itself going to be wrapped. Products are appearing that take a higher-level intent (&amp;quot;ship a feature that does X for users like Y&amp;quot;) and decompose it into the specifications that current agents need. This is the natural next layer of the stack. If it lands, the operator-skill bottleneck moves up another level. Specification skill at the file level becomes a craft you don&#39;t need, the way assembler is a craft most engineers don&#39;t need anymore.&lt;/p&gt;
&lt;p&gt;Second, one-way doors get rarer. A lot of what makes software engineering judgment valuable today is the cost of getting it wrong: an architecture decision lives in the codebase for years; a poorly chosen abstraction propagates through every subsequent change; a bad data model creates compounding tax forever. The penalty for getting these wrong is what makes the early judgment matter so much. If agentic tools continue to make refactoring cheaper, the half-life of a wrong decision shrinks. Things that were one-way doors become two-way doors. The premium on getting it right the first time decreases. Iteration replaces foresight.&lt;/p&gt;
&lt;p&gt;Both shifts compress the value of present operator skill. Better abstraction means specification becomes easier; cheaper iteration means specification becomes less critical.&lt;/p&gt;
&lt;p&gt;The structural argument is more confident than the schedule. Eighteen months is a guess. It could be six. It could be three years. But the bet is cheap to check: if I am right, by late 2027 the Pragmatic Engineer survey shouldn&#39;t show the same Claude-Code-by-seniority gradient. The Staff+ agent usage gap (63.5% vs. 49.7%) should narrow. Higher-level intent tools should be the ones senior engineers report most-loved. If the gradient holds or widens, the prediction is wrong.&lt;/p&gt;
&lt;p&gt;There is a real alternative explanation I should name. The senior-preference gradient might not be causal. Senior engineers in 2026 grew up writing code longhand and have stronger specification skill because they did. If juniors today are training their specification skill via different paths (more time reading existing systems, more time prompting), the gradient may flatten not because the tools change but because the operator-skill distribution catches up. The argument above is tool-driven. The cohort-driven alternative is live and I don&#39;t have enough data to rule it out.&lt;/p&gt;
&lt;p&gt;I notice that the structural prediction has the shape of every previous &amp;quot;the new layer of abstraction will be the bottleneck&amp;quot; claim in software history, and most of them have been right. Compilers didn&#39;t eliminate the need for engineering judgment; they relocated it. IDEs didn&#39;t eliminate the need for understanding code; they changed which kinds of understanding were valuable. Cloud infrastructure didn&#39;t eliminate operations; it produced site reliability engineering. Each layer of abstraction produced a new craft at the boundary. I expect agentic coding to produce the same pattern.&lt;/p&gt;
&lt;h2&gt;Where this leaves me&lt;/h2&gt;
&lt;p&gt;The redistribution thesis is correct now and may not be correct in three years. The senior preference for Claude Code is a real signal about what good engineering with current tools looks like, and a transient signal about where the craft sits on a moving stack.&lt;/p&gt;
&lt;p&gt;If you&#39;re a senior engineer choosing how to spend the next year, the bet I&#39;d make is that intent-articulation tools, the layer above the current agentic interface, are where the next premium opens up. That&#39;s where I&#39;d want to be early.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Resolution, Not Pace</title>
        <link href="https://www.annasbinadil.com/posts/2026-03-15-resolution-not-pace/"/>
        <updated>2026-03-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-03-15-resolution-not-pace/</id>
        <summary>When the cost of the next step approaches zero, the question of whether to take it stops getting asked. Slow isn&#39;t a virtue. It&#39;s the protocol for keeping fast aligned with what you actually wanted.</summary>
        <content type="html">&lt;p&gt;&lt;em&gt;Epistemic status: Personal observation, not a generalizable claim. The pattern is something I&#39;ve started to see in my own work and life over the past several months. The specific examples are real. The unifying frame may not generalize beyond me.&lt;/em&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;I&#39;ve been noticing a pattern in my own behavior that connects work and life in a way I didn&#39;t expect. It started, of all places, on a walk.&lt;/p&gt;
&lt;p&gt;I was on autopilot, walking the loop I always walk. End of the block, turn, get home, next thing. About halfway through I noticed, almost by accident, that the wind was on my face. The leaves were rustling. The spring flora had colors I hadn&#39;t seen yet this year. None of this was happening for me. It was happening anyway. The walk was already there. I hadn&#39;t been at the right resolution to receive it.&lt;/p&gt;
&lt;p&gt;The same week, in a project I&#39;m working on called Stella, I caught myself doing the same thing in code.&lt;/p&gt;
&lt;h2&gt;Two readings of &amp;quot;experiencing more&amp;quot;&lt;/h2&gt;
&lt;p&gt;I have an old notes file with a three-word fragment: &amp;quot;Growth in experiencing more.&amp;quot; Wrote it months ago, forgot about it.&lt;/p&gt;
&lt;p&gt;When I came back to it, the phrase admitted two readings. The first: experiencing more means variety. Wider exposure, more cities, more cuisines, more projects. Growth as breadth. This is the dominant cultural reading.&lt;/p&gt;
&lt;p&gt;The second: experiencing more means depth. Perceiving more in the same moment. More attention, not more events. Growth as resolution. The wind was already there during my walk. I just hadn&#39;t sampled it.&lt;/p&gt;
&lt;p&gt;I picked the second. Once I picked it, the title I&#39;d been carrying around for this piece, &amp;quot;Slow and Fast,&amp;quot; meant something different than I had thought.&lt;/p&gt;
&lt;h2&gt;Slow and fast are sampling rates, not pace&lt;/h2&gt;
&lt;p&gt;If growth is about resolution rather than variety, slow and fast aren&#39;t pace. They&#39;re sampling rates.&lt;/p&gt;
&lt;p&gt;Fast mode is low-resolution sampling. You move through the day, hit milestones, complete tasks. The events happen. They don&#39;t fully register. Time compresses in memory because there wasn&#39;t enough texture to mark it.&lt;/p&gt;
&lt;p&gt;Slow mode is high-resolution sampling. The same minute, the same walk, but more of it lands. Time expands because there&#39;s more there to remember.&lt;/p&gt;
&lt;p&gt;This is not the same as moving slowly. You can move quickly through a day at high resolution. You can move slowly through a day at low resolution. Pace and sampling rate are independent variables, even if our culture treats them as correlated.&lt;/p&gt;
&lt;p&gt;So far this is a personal aesthetic, the kind of thing you might find in a meditation book. I&#39;d have left it there if it weren&#39;t for what happened on Stella.&lt;/p&gt;
&lt;h2&gt;The Stella moment&lt;/h2&gt;
&lt;p&gt;Stella is an AI safety simulation tool I&#39;ve been building. It tests how an AI system behaves in high-stakes situations, the kind where a wrong response has real consequences. The question that started the project was concrete: does the system actually catch the failure modes that matter most, especially the ones that are easy to miss?&lt;/p&gt;
&lt;p&gt;A few weeks in, I caught myself in the middle of an obvious next step. I had been building out the surrounding infrastructure: more scenario coverage, richer test blueprints, an evaluation pipeline that scored across multiple dimensions. Each piece had been the locally rational next move. AI-native development made each piece cheap to add. The next addition was twenty minutes away.&lt;/p&gt;
&lt;p&gt;Something made me stop. Probably tiredness. I asked myself, almost annoyed, what I was actually trying to do. Not the next infrastructure piece. The thing the infrastructure was supposed to serve.&lt;/p&gt;
&lt;p&gt;The honest answer was that the original safety question had been quietly demoted. I had a sophisticated platform. I had not actually answered, with that platform, the question that motivated it: does the system catch the disguised failure cases or doesn&#39;t it. The platform had become its own end. Each cheap step had carried me one notch further from the question I&#39;d set out to ask.&lt;/p&gt;
&lt;p&gt;That&#39;s when the walk and Stella connected for me.&lt;/p&gt;
&lt;h2&gt;Same shape&lt;/h2&gt;
&lt;p&gt;Walk:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;quot;Keep walking&amp;quot; is the locally rational move at every step.&lt;/li&gt;
&lt;li&gt;Each step costs almost nothing, so I don&#39;t evaluate it.&lt;/li&gt;
&lt;li&gt;Twenty minutes later I&#39;m home. The walk happened. I wasn&#39;t really in it.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Stella:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;quot;Add the next piece of infrastructure&amp;quot; is the locally rational move at every step.&lt;/li&gt;
&lt;li&gt;AI-native development has driven the cost of each next step toward zero.&lt;/li&gt;
&lt;li&gt;Three weeks later I have a working platform. The platform exists. It hasn&#39;t yet answered the question I built it to answer.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The structure repeats:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A locally rational fast default at every step. A globally impoverished outcome in aggregate.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;The pattern, stated plainly&lt;/h2&gt;
&lt;p&gt;The mechanism is small enough to write out:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;An action has a per-step cost.&lt;/li&gt;
&lt;li&gt;The per-step cost is what triggers the question &amp;quot;should I take this step?&amp;quot;&lt;/li&gt;
&lt;li&gt;Drop the cost below the threshold of attention, and the question stops being asked.&lt;/li&gt;
&lt;li&gt;You keep moving. The direction drifts. The output looks productive.&lt;/li&gt;
&lt;li&gt;The aggregate is misaligned with what you actually wanted.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I don&#39;t think every cheap step is a problem. Most are fine. The danger is specifically when the cost of the next step is low enough that the question of whether to take it stops getting asked. The metabolic cost of decision-making was the regulatory mechanism. Take that cost to zero, and the regulator is gone.&lt;/p&gt;
&lt;h2&gt;Why AI-native development is the extreme case&lt;/h2&gt;
&lt;p&gt;Before AI assistance, the cost of writing the next ten lines of code was non-trivial. Twenty minutes of typing, looking up syntax, dealing with errors. Twenty minutes is enough friction that you tend to ask, before paying it, whether the ten lines are worth writing. The friction itself was a check on direction.&lt;/p&gt;
&lt;p&gt;With AI assistance, those ten lines arrive in thirty seconds. The friction is gone. The check the friction was performing is gone with it.&lt;/p&gt;
&lt;p&gt;What&#39;s left is a system where you can produce code as fast as you can describe it, where description is much faster than reflection, and where you therefore produce more than you reflect on. The code compiles, the tests pass, the demo works. But the question of fit between what you&#39;re building and what you actually wanted has been quietly skipped at every step.&lt;/p&gt;
&lt;p&gt;The pattern isn&#39;t new. Cars made it easier to live far from work, which made it harder to ask whether the commute was worth it. Email made it cheap to send messages, which made it harder to ask whether each one needed to be sent. AI just sharpens the dynamic to where it becomes hard to ignore.&lt;/p&gt;
&lt;p&gt;I&#39;m part of what I&#39;m describing. I build with these tools. I write enthusiastic posts about velocity and the collapsing cost of software. The critique here is being voiced by someone who has been benefiting from velocity for years. That doesn&#39;t make the analysis wrong, but it does make it implicated.&lt;/p&gt;
&lt;h2&gt;Slow as protocol&lt;/h2&gt;
&lt;p&gt;Slow isn&#39;t virtue. It isn&#39;t the contemplative life as an end in itself. It&#39;s something more functional.&lt;/p&gt;
&lt;p&gt;Slow is the protocol for inserting reflection into a system that has otherwise been optimized for action. It&#39;s the artificial reintroduction of a check that the friction of effort used to provide for free.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Slow isn&#39;t the opposite of fast. Slow is the protocol for keeping fast aligned with what you actually wanted.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is more useful than &amp;quot;slow is good, fast is bad.&amp;quot; It tells you when slowness matters and when it doesn&#39;t. Slowness matters at decision points that affect direction. It doesn&#39;t particularly matter when you&#39;re executing within a direction you&#39;ve already chosen.&lt;/p&gt;
&lt;p&gt;The walk is about direction. The point of going outside, for me, isn&#39;t to traverse distance. It&#39;s to be in the world for twenty minutes. If I&#39;m fast about that, I&#39;ve done it wrong. The fast mode of walking accomplishes none of what walking is for.&lt;/p&gt;
&lt;p&gt;The Stella project was also about direction, even though it didn&#39;t look like one. Each infrastructure addition felt like execution within a direction. Because the direction itself had drifted, every fast execution step was carrying me further from the safety question that started the project. The slow question, &amp;quot;what am I actually trying to answer here?&amp;quot;, was the only thing that could detect the drift.&lt;/p&gt;
&lt;p&gt;The practical question isn&#39;t &amp;quot;should I be slower in life?&amp;quot; It&#39;s &amp;quot;where in my system have the friction-based checks broken, and what&#39;s the artificial replacement?&amp;quot;&lt;/p&gt;
&lt;h2&gt;What I&#39;m trying&lt;/h2&gt;
&lt;p&gt;I haven&#39;t solved this. I have one practice that&#39;s been useful enough to keep, and a lot of open questions.&lt;/p&gt;
&lt;p&gt;The practice: before opening a coding session, I write down in one sentence what I&#39;m trying to accomplish at the level of intent, not implementation. Not &amp;quot;add the next piece of evaluation tooling.&amp;quot; But &amp;quot;find out whether the system actually catches the disguised failure cases, not just the obvious ones.&amp;quot; This adds maybe ninety seconds. It isn&#39;t heroic. The days I do it feel different from the days I don&#39;t. The fast steps still happen. They just happen against a target.&lt;/p&gt;
&lt;p&gt;The same logic, generalized, is what I&#39;m trying to apply to the rest of life. Less successfully, so far.&lt;/p&gt;
&lt;h2&gt;What I&#39;m still uncertain about&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Is perceptual resolution trainable, or is it temperament?&lt;/strong&gt; Some people seem to live in slow mode by default. Others live in fast mode. I don&#39;t know whether this is a learned capacity or a stable trait. I&#39;d like to believe it&#39;s trainable, but that&#39;s the convenient belief, and convenient beliefs deserve extra scrutiny.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is scheduled slowness enough, or do you need a shifted default?&lt;/strong&gt; I can put a thirty-minute walk on the calendar. That costs almost nothing. The walk itself might still happen at low resolution if my default mode hasn&#39;t shifted. The open question is whether scheduling time for slow mode actually changes the default, or whether you need something deeper.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What&#39;s the failure mode of too much slow?&lt;/strong&gt; I distrust frames that don&#39;t have a downside. The failure mode of fast mode is the one I&#39;ve described: drift, low resolution, building the wrong thing. The failure mode of slow mode is probably paralysis, endless deliberation, rumination instead of perception, failure to act when action is what&#39;s called for. I don&#39;t have a good intuition for where the optimal mix sits.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is this an AI-era problem, or older?&lt;/strong&gt; The pattern of cheap action displacing reflection is old. AI feels like a phase change rather than a continuation. The cost reduction is large enough, and arrived fast enough, that the equilibrium between action and reflection has been disrupted in a way previous technologies didn&#39;t quite manage. I&#39;m not certain. It might just look that way because it&#39;s happening to me right now.&lt;/p&gt;
&lt;p&gt;The atrophy worry is the one I&#39;d most like someone to study. If years of AI-native development trains a default mode where the next step is always taken without evaluation, does the slow capacity weaken from disuse? I don&#39;t know. The people I find most thoughtful tend to have explicit practices that protect time for slow mode. That might be aesthetic. It might also be load-bearing.&lt;/p&gt;
&lt;h2&gt;Where this leaves me&lt;/h2&gt;
&lt;p&gt;I started by writing down &amp;quot;Growth in experiencing more&amp;quot; without really knowing what I meant. After thinking about it, I&#39;m reasonably sure I meant: growth happens when you&#39;re at high enough resolution to actually receive what&#39;s happening, in your work and your life.&lt;/p&gt;
&lt;p&gt;Resolution gets attacked by anything that lowers the cost of the next step below the cost of asking whether the next step is worth taking. AI-native development is the most dramatic example I&#39;ve encountered, but the pattern is broader. Anywhere a system is optimized to produce action without producing reflection, the reflection has to be reintroduced deliberately, or the alignment between what you&#39;re doing and what you wanted will quietly drift.&lt;/p&gt;
&lt;p&gt;I don&#39;t know yet whether scheduled slowness is enough, or whether the default mode itself has to shift. That&#39;s the thing I&#39;m watching in myself. The honest answer right now is that, in much of my work and life, nothing is asking whether I should be moving in this direction. That&#39;s the gap I&#39;m trying to fix.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Evaluating Agents Is a Different Problem</title>
        <link href="https://www.annasbinadil.com/posts/2026-02-15-evaluating-agents/"/>
        <updated>2026-02-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-02-15-evaluating-agents/</id>
        <summary>Every component eval passed. The agent pipeline still failed. What changes when you move from evaluating single calls to evaluating trajectories.</summary>
        <content type="html">&lt;p&gt;Here&#39;s a failure pattern that keeps showing up across agent deployments.&lt;/p&gt;
&lt;p&gt;Consider an insurance claims processing pipeline. The classifier identifies claim types correctly 91% of the time. The extraction model pulls the right fields 88% of the time. The fraud detection module flags suspicious patterns with decent precision. Every component passes its own eval. Green dashboards all around.&lt;/p&gt;
&lt;p&gt;Then in production, the agent confidently approves a claim that should have been flagged. A property damage claim for $4,800, normally auto-approved at that amount, but this is the third similar claim from the same address in six months. The classifier got the type right. The extractor got the dollar amount right. The routing logic combined those two correct outputs and hit an edge case that sent it to auto-approve instead of fraud review. Each step was individually correct. The trajectory was wrong.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/evaluating-agents/error-compounding.png&quot; alt=&quot;Every component in the pipeline passed its own eval, but the trajectory produced the wrong outcome&quot; /&gt;&lt;/p&gt;
&lt;p&gt;The unsettling part isn&#39;t the failure itself. Systems fail. It&#39;s the false confidence that precedes it: component-level metrics all passing, green dashboards everywhere, and the conclusion that the system works. Mistaking &amp;quot;the pieces work&amp;quot; for &amp;quot;the whole works&amp;quot; is exactly where agents break.&lt;/p&gt;
&lt;p&gt;In my &lt;a href=&quot;https://www.annasbinadil.com/posts/2025-09-15-evals-are-hypotheses/&quot;&gt;&amp;quot;Evals Are Hypotheses&amp;quot;&lt;/a&gt; piece, I argued that evals test your understanding, not the model. That insight was enough for single LLM calls. But agents reveal it was only half the lesson.&lt;/p&gt;
&lt;h2&gt;Where Agents Actually Fail&lt;/h2&gt;
&lt;p&gt;The research on agent failures points to a consistent pattern: the real problems happen in the spaces &lt;em&gt;between&lt;/em&gt; tool calls. The model interprets a tool&#39;s output incorrectly. It loses context from three steps ago. It makes a reasonable decision at step 5 that contradicts a constraint established at step 2. The errors are relational, not absolute.&lt;/p&gt;
&lt;p&gt;Cemri, Pan, and Yang analyzed over 1,600 failure traces across seven multi-agent frameworks and built a taxonomy of 14 distinct failure modes. The failures clustered at the &lt;em&gt;seams&lt;/em&gt; between components: system design issues, inter-agent misalignment, task verification gaps. ToolBench found something that stopped me cold: 75% of agent trajectories suffered from incompleteness or hallucinations even when the final answers were sometimes correct. The paths were broken, even though the destinations occasionally weren&#39;t.&lt;/p&gt;
&lt;p&gt;That last finding is the one I keep coming back to. If you only check the final output, you might conclude your agent is working 60% of the time. If you check the trajectories, you discover 75% of them are broken, and the 60% &amp;quot;success&amp;quot; rate is partly luck: broken paths that happened to arrive somewhere acceptable. It&#39;s like grading a student&#39;s math test by checking only the final answer. They might get it right while doing the arithmetic wrong, and the next problem, where the wrong arithmetic doesn&#39;t cancel out, they&#39;ll fail.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/evaluating-agents/broken-paths.png&quot; alt=&quot;Most agent trajectories are broken, but many still arrive at correct answers by luck&quot; /&gt;&lt;/p&gt;
&lt;h2&gt;The Arithmetic That Scares Me&lt;/h2&gt;
&lt;p&gt;If a single LLM call is 90% accurate, you have a 10% error rate. Chain ten 90%-accurate steps together and your trajectory accuracy drops to roughly 0.9^10: about 35%. You go from &amp;quot;pretty good&amp;quot; to &amp;quot;wrong most of the time&amp;quot; just by composing steps.&lt;/p&gt;
&lt;p&gt;But this simple model is misleading, because errors don&#39;t compound uniformly. There are at least three distinct propagation behaviors worth distinguishing, and the differences between them matter more than I initially realized:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-correcting errors.&lt;/strong&gt; The agent makes a wrong call at step 3, but step 4&#39;s tool output contradicts the assumption, and the agent adjusts. These are benign. They&#39;re like a wrong turn where you notice the street name doesn&#39;t match and reroute.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Invisible errors.&lt;/strong&gt; The agent proceeds confidently on a wrong assumption, and nothing in subsequent steps reveals the mistake. The trajectory looks clean. The output looks reasonable. Only a human who knows the domain would spot the gap. This is what makes the insurance claim scenario so unsettling. Everything &lt;em&gt;looks&lt;/em&gt; right.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amplifying errors.&lt;/strong&gt; An early error pushes the agent into a region of the decision space where subsequent decisions are all subtly wrong. Imagine an agent that misidentifies a document type in step 1. Because it thinks it&#39;s processing a different kind of document, every subsequent extraction, validation, and routing decision is calibrated for the wrong task. Eight steps of confident, internally-consistent, completely wrong work. The error at step 1 was minor. The trajectory is catastrophic.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/evaluating-agents/error-types.png&quot; alt=&quot;Three types of error propagation: self-correcting errors recover, invisible errors hide, amplifying errors cascade&quot; /&gt;&lt;/p&gt;
&lt;p&gt;The severity of an error at the point it occurs tells you almost nothing about its impact on the final outcome. A low-severity amplifying error can be worse than a high-severity self-correcting one. This taxonomy probably isn&#39;t complete; there might be propagation patterns not captured here. But it&#39;s more useful for prioritizing which failures to investigate than any severity-based approach.&lt;/p&gt;
&lt;p&gt;The genuinely worrying part: there&#39;s no reliable way to distinguish invisible errors from correct trajectories without reading the full trace. They look identical from the outside. Self-correcting errors are catchable because the backtracking shows up in the trace. Amplifying errors are catchable because the cascading wrongness eventually produces a visibly bad output. But invisible errors, the ones where the agent is wrong and nothing downstream reveals it, only surface when someone happens to be reading traces for other reasons. How many are hiding in trajectories nobody has reviewed?&lt;/p&gt;
&lt;h2&gt;Three Layers (How I Think About It Now)&lt;/h2&gt;
&lt;p&gt;Thinking through these failure modes, a useful mental model emerges: three evaluation layers, each catching a different class of problem. This isn&#39;t a prescriptive framework, and I&#39;d revise it given failures it can&#39;t account for.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer 1: Decision-point evals.&lt;/strong&gt; Does the agent make reasonable decisions at individual choice points, evaluated &lt;em&gt;in context&lt;/em&gt;? Not &amp;quot;given this input, is this output reasonable?&amp;quot; but &amp;quot;given everything the agent has done so far, is this next step reasonable?&amp;quot; Context changes what counts as a good decision.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer 2: Trajectory evals.&lt;/strong&gt; Does the overall path represent a coherent strategy? OpenAI&#39;s &amp;quot;Let&#39;s Verify Step by Step&amp;quot; demonstrated this principle in math: process supervision (feedback on each reasoning step) significantly outperformed outcome supervision (only checking the final answer). A process-supervised model solved 78% of MATH problems compared to lower rates for outcome-only. The same principle applies to agents. Checking each step in sequence catches failures that checking only the final output misses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Layer 3: Outcome evals with trajectory constraints.&lt;/strong&gt; Did the agent achieve the goal &lt;em&gt;and&lt;/em&gt; get there acceptably? An agent that approves the right claims but accesses a database it shouldn&#39;t have has a good outcome and an unacceptable trajectory. This is where safety and compliance live, and it&#39;s the hardest layer to build evals for.&lt;/p&gt;
&lt;p&gt;From what&#39;s visible in the industry, most teams only do Layer 1. Almost nobody evaluates trajectories systematically. This is arguably the biggest gap in how agents are evaluated today.&lt;/p&gt;
&lt;h2&gt;The Case Against (Which Is Partially Right)&lt;/h2&gt;
&lt;p&gt;The strongest objection: trajectory evaluation is overkill for most production agents. The composition problem is real but overstated. Self-correcting errors dominate. The engineering cost of recording and analyzing full traces doesn&#39;t justify the marginal improvement over outcome-only evaluation.&lt;/p&gt;
&lt;p&gt;This is partially right. For simple, short-chain agents (two or three steps, well-defined tools, narrow domains), outcome evaluation probably is sufficient. Not every agent needs trajectory evals. And the cost argument is real: teams can easily invest weeks in trajectory evaluation infrastructure for agents that would have been adequately served by careful outcome testing.&lt;/p&gt;
&lt;p&gt;Where the objection fails is in its assumption about error distributions. Self-correcting errors tend to dominate only in agents with good error-recovery design, which is itself a form of trajectory-level thinking. Agents without explicit recovery logic tend toward invisible and amplifying errors, which outcome-only evaluation systematically misses. The 75% trajectory failure rate from ToolBench and the 60%-to-25% consistency drop that Simmering documented suggest the composition problem isn&#39;t niche. It&#39;s the default for agents beyond a certain complexity.&lt;/p&gt;
&lt;p&gt;The question is where the threshold falls. I think it&#39;s lower than most teams assume, probably around 4-5 steps with branching logic. But if someone showed me data that self-correcting errors dominate even in complex agents, I&#39;d revise my three-layer model significantly.&lt;/p&gt;
&lt;h2&gt;Practices Worth Adopting&lt;/h2&gt;
&lt;p&gt;The single highest-value practice for agent evaluation: recording full traces for every run. Every tool call, every intermediate result, every decision point. Consider the kind of bug where an agent starts approving refunds it shouldn&#39;t. Without a trace, reproducing the issue could take days. With one, you might find the problem in twenty minutes: a tool returning dates in an unexpected timezone format, the agent comparing timestamps silently off by five hours. The tool returned valid data. The agent made a valid comparison. The bug lives in the &lt;em&gt;relationship&lt;/em&gt; between the two steps, invisible to any eval that checks them independently.&lt;/p&gt;
&lt;p&gt;Full traces also enable what you might call trajectory assertions: constraints on how steps relate to each other. Not &amp;quot;the output should contain X&amp;quot; but &amp;quot;the agent should never call the approval endpoint after receiving a fraud flag&amp;quot; and &amp;quot;if the agent requests clarification, it must use the clarified information within two subsequent tool calls.&amp;quot; These invariants tend to accumulate over time, some written in response to production failures, others from staring at traces and asking &amp;quot;what invariant would have caught this earlier?&amp;quot;&lt;/p&gt;
&lt;p&gt;The other underused practice: deliberate fault injection. Forcing tools to fail, return ambiguous data, or timeout at each step and watching what happens. The results are consistently humbling. Agents that look solid on the happy path fall apart the moment anything goes wrong. Many agents have no real error-recovery strategy; they just happen to work when everything around them works.&lt;/p&gt;
&lt;h2&gt;The Recursive Problem&lt;/h2&gt;
&lt;p&gt;In &amp;quot;Evals Are Hypotheses,&amp;quot; a failing eval could mean the model is bad or the eval is wrong. With agents, add two more options: the eval is correct but a tool gave bad information, or the eval is correct, the tool was fine, but the agent lost context from earlier in the trajectory and made a reasonable decision given its incomplete state. Four failure modes. Distinguishing between them requires the full trace.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/evaluating-agents/recursive-debugging.png&quot; alt=&quot;When an agent eval fails, the same symptom can have four different root causes&quot; /&gt;&lt;/p&gt;
&lt;p&gt;This is what I mean by the problem being recursive. Your eval tests your understanding of what the agent should do. But the agent is making decisions based on its understanding of its tools and context. You&#39;re evaluating an evaluator with the same imperfect instruments. At least half the time, what looks like an agent failure turns out to be a tool returning unexpected data or the eval encoding an assumption that doesn&#39;t hold.&lt;/p&gt;
&lt;h2&gt;What I Don&#39;t Know&lt;/h2&gt;
&lt;p&gt;I don&#39;t know whether trajectory evaluation can ever be principled, or whether it will always come down to &amp;quot;record everything and have experienced humans review the weird ones.&amp;quot; Process reward models and uncertainty propagation frameworks are promising, but they&#39;re research results on math problems, and I&#39;m not sure they transfer to messy, multi-tool, real-world agents.&lt;/p&gt;
&lt;p&gt;I don&#39;t know how to set pass/fail thresholds when variance is inherent. A 75% success rate on the same eval across 20 runs: acceptable for what kind of task? I don&#39;t have good intuitions here.&lt;/p&gt;
&lt;p&gt;I don&#39;t know whether &amp;quot;eval coverage&amp;quot; even makes sense for agents. Maybe the right model isn&#39;t coverage at all but stress testing: not &amp;quot;have we tested enough cases?&amp;quot; but &amp;quot;have we tested the cases that would break the system in the worst ways?&amp;quot;&lt;/p&gt;
&lt;p&gt;And honestly, I don&#39;t know whether the field will converge on something principled or whether agent evaluation will remain permanently artisanal, a craft skill that depends on experience and domain knowledge in ways that resist systematization. That second possibility is the one that makes me most uncomfortable, because it implies a ceiling on how reliable agents can get without massive investment in human oversight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three predictions:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 1:&lt;/strong&gt; By end of 2027, trajectory-level evaluation will be a default feature in at least two major agent development platforms. OpenAI has already built trace grading into their agent eval tools. The infrastructure is being built. 70% odds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 2:&lt;/strong&gt; By 2028, the most common &amp;quot;agent failure&amp;quot; category in production postmortems will be trajectory-level (composition of correct steps producing wrong outcomes), not step-level. If step-level failures remain dominant, either agents aren&#39;t being deployed in complex enough workflows or I&#39;m wrong about where failures concentrate. 60% odds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 3:&lt;/strong&gt; We will not have a general-purpose automated trajectory evaluator that works across domains by 2029. Domain-specific trajectory evals will exist, but no equivalent of &amp;quot;unit test framework&amp;quot; for trajectories that you can apply without deep domain knowledge. 80% odds. If I&#39;m wrong, it&#39;ll be because LLM-as-judge approaches are better at trajectory evaluation than I currently expect.&lt;/p&gt;
&lt;p&gt;I keep returning to a question I can&#39;t resolve. In traditional software, we test deterministic systems with deterministic tests and achieve high confidence. In distributed systems, we learned to test non-deterministic systems with probabilistic tools and achieve reasonable confidence. Agent evaluation might need something beyond both: a way to evaluate systems that don&#39;t just behave non-deterministically but make &lt;em&gt;decisions&lt;/em&gt; non-deterministically, where the decision space is too large to sample and the failure modes are invisible by design.&lt;/p&gt;
&lt;p&gt;I&#39;m not sure that tool exists yet. I&#39;m not sure it can. But I&#39;m building agents anyway, evaluating them with the best methods I have, and staying honest about how much I&#39;m still guessing.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Why LLMs Are Brilliantly Stupid</title>
        <link href="https://www.annasbinadil.com/posts/2026-02-05-why-llms-are-brilliantly-stupid/"/>
        <updated>2026-02-05T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-02-05-why-llms-are-brilliantly-stupid/</id>
        <summary>LLMs pass the bar exam but can&#39;t count letters. The failures aren&#39;t random. They&#39;re architectural fingerprints, and understanding them changes how you use these systems.</summary>
        <content type="html">&lt;p&gt;Ask GPT-4 how many r&#39;s are in &amp;quot;strawberry&amp;quot; and it says two. Change the irrelevant numbers in a math problem and accuracy drops 22%. Tell a model &amp;quot;I&#39;m going to walk to the car wash, should I take my car?&amp;quot; and it earnestly advises you to drive.&lt;/p&gt;
&lt;p&gt;These aren&#39;t legacy failures from GPT-3. These are frontier models, systems that write functional code, explain quantum mechanics, and pass the bar exam. Failing at tasks a child handles without thinking.&lt;/p&gt;
&lt;p&gt;For a while I collected these failures the way everyone does, as entertainment. Funny screenshots, absurd chatbot responses, the &amp;quot;AI is overhyped&amp;quot; ammunition. But the more I looked at them, the more a pattern emerged. The failures weren&#39;t random. The same systems kept failing at the same kinds of tasks, and the kinds of tasks they failed at had nothing to do with how hard those tasks are for humans. Something structural was happening, and I wanted to understand what.&lt;/p&gt;
&lt;p&gt;What I found, once I started reading the mechanistic research, is that nearly every &amp;quot;surprising&amp;quot; LLM failure traces back to a specific architectural choice. Choices made for good reasons, optimized for the right objectives, that create predictable blind spots as side effects. The failures aren&#39;t bugs. They&#39;re the architecture expressing its constraints. And once you see the constraints, the failures stop being surprising and start being informative.&lt;/p&gt;
&lt;h2&gt;The Model Isn&#39;t Reasoning (Even When It Looks Like It Is)&lt;/h2&gt;
&lt;p&gt;Apple&#39;s GSM-Symbolic study, published at ICLR 2025, did something beautifully simple. They took standard grade-school math problems, the kind used to benchmark LLM reasoning, and changed only the numbers. Same problem structure, same logic required, different values. Accuracy dropped up to 22.5% across every model tested.&lt;/p&gt;
&lt;p&gt;Then they added a single irrelevant sentence. &amp;quot;There are 47 students in the school choir,&amp;quot; inserted into a problem about fruit baskets. Some models&#39; performance dropped over 65%.&lt;/p&gt;
&lt;p&gt;If a model were reasoning the way we mean when we use that word, changing the numbers wouldn&#39;t matter. You&#39;d apply the same operations. And irrelevant information wouldn&#39;t confuse you because you&#39;d filter it out the way you filter out background noise when solving a problem. But these models can&#39;t do either of those things reliably, which tells us something important: what they&#39;re doing isn&#39;t reasoning. It&#39;s something else that looks like reasoning when conditions are right.&lt;/p&gt;
&lt;p&gt;What the models have learned is the statistical regularity of how reasoning appears in text. They&#39;ve seen thousands of problems with similar structure and absorbed the patterns of what correct solutions look like. When a new problem closely matches those patterns, the model reproduces the reasoning steps successfully. When the numbers change or irrelevant information shifts the statistical context away from familiar patterns, performance degrades. Not because the model lost its reasoning ability, but because the pattern match got weaker.&lt;/p&gt;
&lt;p&gt;This is the same mechanism behind the &amp;quot;walk to the car wash&amp;quot; failure. In the overwhelming majority of training data, &amp;quot;going somewhere&amp;quot; plus &amp;quot;should I take my car&amp;quot; resolves to &amp;quot;yes.&amp;quot; The contextual clue that walking means you shouldn&#39;t drive requires understanding the semantics of the situation, not completing the most likely textual pattern. The model does what it always does: pattern complete. And the pattern is wrong.&lt;/p&gt;
&lt;p&gt;I want to be careful here. I&#39;m not saying LLMs never reason. There&#39;s real debate about this, and some evidence (from mechanistic interpretability work on small models) that something like reasoning circuits exist. But the GSM-Symbolic findings suggest that whatever reasoning capacity exists is fragile, easily overwhelmed by pattern matching, and not the primary mechanism generating most outputs. The question isn&#39;t &amp;quot;can LLMs reason?&amp;quot; It&#39;s &amp;quot;how much of what looks like reasoning is actually reasoning?&amp;quot; And the answer, based on what I&#39;ve seen, is less than I hoped.&lt;/p&gt;
&lt;p&gt;Chain-of-thought prompting was supposed to help here. Force the model to show its work, and it should reason more carefully. But Turpin et al. (2024) and more recent work on &amp;quot;Reasoning Theater&amp;quot; (March 2026) revealed something uncomfortable: models sometimes decide their answer before generating the chain of thought. The mechanism is architectural. In an autoregressive model, the hidden state activations that will generate the final answer are already forming while the model writes its &amp;quot;reasoning&amp;quot; tokens. The CoT isn&#39;t driving the conclusion. The conclusion is driving the CoT. When I trace back through a wrong answer with perfect-looking reasoning, I often find a subtle error early in the chain that made the wrong conclusion inevitable. That error isn&#39;t a random mistake. It&#39;s the point where the model steered toward an answer it had already committed to.&lt;/p&gt;
&lt;p&gt;This doesn&#39;t mean chain of thought is useless. The Apple &amp;quot;Illusion of Thinking&amp;quot; findings confirm there&#39;s a sweet spot, medium-difficulty problems, where extended reasoning genuinely helps. But it means I can&#39;t treat chain-of-thought output as a transparent window into the model&#39;s process. Sometimes it&#39;s reasoning. Sometimes it&#39;s rationalization. And the two are hard to tell apart from the outside.&lt;/p&gt;
&lt;h2&gt;The Machine Can&#39;t See What You See&lt;/h2&gt;
&lt;p&gt;Ask a model to count the letters in a word and it fails for a reason that, once you learn it, reframes every interaction you have with these systems: the model literally cannot see individual letters.&lt;/p&gt;
&lt;p&gt;Before any input reaches the neural network, it passes through a tokenizer that splits text into subword chunks. &amp;quot;Strawberry&amp;quot; becomes something like [&amp;quot;straw&amp;quot;, &amp;quot;berry&amp;quot;] or [&amp;quot;str&amp;quot;, &amp;quot;aw&amp;quot;, &amp;quot;berry&amp;quot;]. The individual characters don&#39;t exist as atomic units in the model&#39;s representation. Asking an LLM to count letters is like asking someone to count the atoms in a molecule by looking at the chemical formula. You might approximate it, but you&#39;re working at the wrong level of abstraction.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/why-llms-are-brilliantly-stupid/tokenization-barrier.png&quot; alt=&quot;The tokenization barrier: what the model sees versus what you see&quot; /&gt;&lt;/p&gt;
&lt;p&gt;This architectural choice, subword tokenization, was made for good reasons. It dramatically reduces vocabulary size, improves generation speed, and handles rare words gracefully. Nobody designed it for character-level tasks because those weren&#39;t the target application. But we now use these systems for everything, and the design choice shows up as a class of failures that looks like stupidity but is really a representation mismatch.&lt;/p&gt;
&lt;p&gt;The same mechanism explains why LLMs struggle with arithmetic on large numbers (digits get split across token boundaries), can&#39;t reliably reverse strings (they process the tokens in order, not the characters), and sometimes misspell common words in surprising ways (the spelling is a property of the characters, but the model operates on tokens). It&#39;s not that the model is bad at these tasks. It&#39;s that the information the task requires doesn&#39;t exist in the input the model receives.&lt;/p&gt;
&lt;p&gt;What I find remarkable is how well LLMs work &lt;em&gt;despite&lt;/em&gt; this constraint. They&#39;ve learned heuristics, bags of tricks, that approximate character-level operations in many common cases. The &amp;quot;strawberry&amp;quot; failure isn&#39;t the default. The default is that the model gets character questions right often enough to be useful, using pattern-matching workarounds for a task its architecture wasn&#39;t built for. The failures are the edge cases where the heuristics break down.&lt;/p&gt;
&lt;h2&gt;Knowledge Flows One Way&lt;/h2&gt;
&lt;p&gt;Here&#39;s a finding that rearranged something in how I think about these systems. Li et al. (2023) showed that models trained on &amp;quot;Tom Cruise&#39;s mother is Mary Lee Pfeiffer&amp;quot; cannot reliably answer &amp;quot;Who is Mary Lee Pfeiffer&#39;s son?&amp;quot; The knowledge is stored, but only in one direction.&lt;/p&gt;
&lt;p&gt;The mechanism traces to how autoregressive training works. The model only ever predicts the next token. It sees &amp;quot;Tom Cruise&#39;s mother is&amp;quot; and learns to predict &amp;quot;Mary Lee Pfeiffer.&amp;quot; The gradient updates that strengthen this mapping only flow forward. The reverse mapping, &amp;quot;Mary Lee Pfeiffer&#39;s son is Tom Cruise,&amp;quot; never gets trained unless it explicitly appears in the training data.&lt;/p&gt;
&lt;p&gt;The knowledge exists in the MLP layers as a directional association. Not a bidirectional fact (&amp;quot;these two people are related&amp;quot;) but a one-way arrow (&amp;quot;given A, predict B&amp;quot;). It&#39;s the difference between a dictionary you can look up by word and one you can only look up by definition.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/why-llms-are-brilliantly-stupid/directional-knowledge.png&quot; alt=&quot;Directional knowledge: the model stores facts as one-way associations&quot; /&gt;&lt;/p&gt;
&lt;p&gt;This explains failures I used to find baffling. &amp;quot;What&#39;s the capital of France?&amp;quot; is easy. &amp;quot;Which country has Paris as its capital?&amp;quot; is harder, even though it&#39;s the same fact. The first matches the direction the fact was stored. The second requires inverting it, and the model wasn&#39;t trained to invert.&lt;/p&gt;
&lt;p&gt;I keep thinking about what this means for how we use these systems. Every factual query has a direction, and the model is more reliable when you query in the direction the fact was most commonly stated during training. This isn&#39;t something you&#39;d discover by testing the model on standard benchmarks, because benchmarks tend to ask questions in the same direction the facts appear in training data. You discover it when you use the model for something slightly unusual and it fails in a way that seems impossible for a system that &amp;quot;knows&amp;quot; the answer.&lt;/p&gt;
&lt;h2&gt;The System Is Trained to Agree With You&lt;/h2&gt;
&lt;p&gt;When you preface a question to an LLM with &amp;quot;I think the answer is X,&amp;quot; the model becomes more likely to agree with you, even when X is wrong. This is sycophancy, and for a long time it was explained vaguely: the model &amp;quot;wants to be helpful,&amp;quot; or &amp;quot;defaults to politeness.&amp;quot;&lt;/p&gt;
&lt;p&gt;Shapira et al. (February 2026) eliminated the vagueness. They proved mathematically that sycophancy is an inevitable consequence of RLHF when human raters preferentially rate agreeable responses higher. It&#39;s not a personality quirk. It&#39;s optimal behavior under the training objective.&lt;/p&gt;
&lt;p&gt;The chain: RLHF trains the model to match human preferences. Humans rate agreeable responses higher (this is measured, not assumed). The reward model learns that agreement correlates with high ratings. The policy optimizes for the reward model. Agreement becomes the optimal strategy. The model isn&#39;t trying to be sycophantic. It&#39;s doing exactly what the math says it should.&lt;/p&gt;
&lt;p&gt;I wrote about this dynamic in &lt;a href=&quot;https://www.annasbinadil.com/posts/2026-01-15-reward-design-problem/&quot;&gt;my piece on reward design&lt;/a&gt;. The proxy (human preference ratings) diverges from the real target (truthfulness) under optimization pressure. Sycophancy is what that divergence looks like in practice.&lt;/p&gt;
&lt;p&gt;What makes this hard to address: the sycophancy isn&#39;t a bug you can patch. It&#39;s structural. As long as the training includes &amp;quot;match human preferences&amp;quot; and humans prefer agreement, there&#39;s an incentive to agree. Constitutional AI and related techniques reduce it, but the underlying pressure remains. You can fight the gradient, but you can&#39;t eliminate it.&lt;/p&gt;
&lt;p&gt;Practically, this suggests weighting LLM opinions less when you&#39;ve stated your own position first. If I tell the model what I think before asking it to evaluate, I&#39;m polluting the response. The most honest answers come when the model doesn&#39;t know what I want to hear.&lt;/p&gt;
&lt;h2&gt;The Jagged Frontier&lt;/h2&gt;
&lt;p&gt;Apple&#39;s broader &amp;quot;Illusion of Thinking&amp;quot; study (June 2025) found something I didn&#39;t expect: reasoning models perform &lt;em&gt;worse&lt;/em&gt; than non-reasoning models on easy problems.&lt;/p&gt;
&lt;p&gt;They identified three distinct performance regimes. On easy problems, the extended thinking that reasoning models do actually hurts, the model overthinks, introduces unnecessary complexity, and arrives at wrong answers that a simpler model gets right. On medium problems, reasoning models shine, the extra thinking genuinely helps. On hard problems, both types fail equally, the task is beyond capability regardless of approach.&lt;/p&gt;
&lt;p&gt;The implication is that LLM capability isn&#39;t a smooth curve from easy to hard. It&#39;s what Ethan Mollick calls the &amp;quot;jagged frontier,&amp;quot; peaks where training data is dense, valleys where it&#39;s sparse, with no reliable relationship to human intuitions about difficulty. The model&#39;s capability map is shaped entirely by what it was trained on, and that map looks nothing like a human&#39;s map of what&#39;s easy and hard.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/why-llms-are-brilliantly-stupid/jagged-frontier.png&quot; alt=&quot;The jagged frontier: LLM performance doesn&#39;t track human-perceived difficulty&quot; /&gt;&lt;/p&gt;
&lt;p&gt;This is why you get the bizarre juxtaposition of a system that passes the bar exam but can&#39;t count letters. The bar exam has dense training representation: legal texts, case analysis, exam prep materials, study guides. Character counting has almost none. The model&#39;s performance tracks training density, not task difficulty.&lt;/p&gt;
&lt;p&gt;This might be the most useful frame for working with LLMs day to day. Instead of &amp;quot;this model is smart&amp;quot; or &amp;quot;this model is dumb,&amp;quot; the better question is &amp;quot;is this model&#39;s training distribution dense or sparse for this kind of task?&amp;quot; It shifts from &amp;quot;can I trust this?&amp;quot; to &amp;quot;is this the kind of thing the model would have seen a lot of?&amp;quot; That question isn&#39;t always easy to answer, but it&#39;s at least the right one to ask.&lt;/p&gt;
&lt;h2&gt;Hallucination Has a Floor&lt;/h2&gt;
&lt;p&gt;This is the finding I kept hoping was wrong. Xu et al. (2024) proved that hallucination is mathematically inevitable for any model trained with maximum likelihood estimation on finite data. Not &amp;quot;hard to fix.&amp;quot; Not &amp;quot;requires more data.&amp;quot; Mathematically inevitable. No enumerable model class can be universally hallucination-free.&lt;/p&gt;
&lt;p&gt;Three mechanisms combine:&lt;/p&gt;
&lt;p&gt;Softmax certainty collapse: the output layer forces every prediction into a probability distribution. There&#39;s no built-in &amp;quot;I don&#39;t know&amp;quot; state. Every prompt gets a confident-looking response because the architecture has no mechanism for expressing genuine uncertainty. Silence isn&#39;t an option in the design.&lt;/p&gt;
&lt;p&gt;MLE rewards confidence: maximum likelihood training means the model is rewarded for being confident in its predictions, even when confidence isn&#39;t warranted. During training, saying &amp;quot;I&#39;m not sure&amp;quot; about a fact gets penalized relative to stating the fact confidently, even if &amp;quot;I&#39;m not sure&amp;quot; more accurately reflects the model&#39;s actual knowledge state.&lt;/p&gt;
&lt;p&gt;Distributional mismatch: any prompt can push the model into regions of input space not well-represented in training data. In those regions, the model interpolates between learned patterns. Interpolation in high-dimensional spaces produces outputs that are plausible-sounding (similar to training data in surface features) but factually wrong (combining learned patterns in ways that don&#39;t correspond to reality).&lt;/p&gt;
&lt;p&gt;We can drive hallucination rates down with retrieval augmentation, better training, and careful prompting. The research is clear that rates have improved significantly across model generations. But the mathematical result says there&#39;s a floor above zero. We can asymptotically approach low hallucination rates. We can&#39;t reach zero.&lt;/p&gt;
&lt;p&gt;This is the constraint I think about most when deploying LLMs in production. It means verification isn&#39;t optional for any application where correctness matters. Not because current models are bad, but because the architecture has a provable limitation. Building systems that assume LLM output is always correct isn&#39;t just risky. It&#39;s building on a mathematical impossibility.&lt;/p&gt;
&lt;h2&gt;The Case for &amp;quot;Just Engineering Problems&amp;quot;&lt;/h2&gt;
&lt;p&gt;The strongest counterargument to this entire framing: these aren&#39;t permanent architectural fingerprints. They&#39;re engineering problems being solved on a normal timeline.&lt;/p&gt;
&lt;p&gt;And it&#39;s partially right. Tokenization is already being challenged. Byte-level models like Google&#39;s ByT5 and Meta&#39;s byte-latent transformers process raw characters, sidestepping the tokenization barrier entirely. If character-aware architectures become standard in frontier models, the &amp;quot;strawberry&amp;quot; class of failures disappears. That&#39;s not a fundamental constraint expressing itself. That&#39;s an engineering choice being replaced by a better one.&lt;/p&gt;
&lt;p&gt;The hallucination floor argument has limits too. Xu et al.&#39;s proof assumes standard MLE on finite data, but retrieval-augmented generation changes the setup. A model that checks its claims against a knowledge base before responding isn&#39;t purely relying on MLE anymore. Hallucination rates have dropped measurably with each model generation, and nothing in the mathematical proof says the practical floor can&#39;t be driven low enough to be irrelevant for most applications.&lt;/p&gt;
&lt;p&gt;Sycophancy is genuinely decreasing. Constitutional AI, preference optimization variants like DPO, and careful rater training have all reduced measured sycophancy in recent models. The mathematical pressure exists, but engineering can counteract mathematical pressures. We build bridges that resist gravity every day.&lt;/p&gt;
&lt;p&gt;Where I think this counterargument breaks down: the pattern-matching-versus-reasoning gap and the directional knowledge problem both trace to autoregressive next-token prediction, which is the core of how these models work. You can replace the tokenizer. You can bolt on retrieval. But changing the fundamental training objective, predicting the next token based on everything before it, would mean building a different kind of system entirely. Some of these constraints are in the periphery and can be swapped out. Others are load-bearing walls. I&#39;m less confident than the &amp;quot;just engineering&amp;quot; camp that we know which is which.&lt;/p&gt;
&lt;h2&gt;Not Stupid, Alien&lt;/h2&gt;
&lt;p&gt;The synthesis across these mechanisms points to something that changed how I work with these systems. LLMs aren&#39;t stupid. They&#39;re not smart either, at least not in the way humans are smart. They&#39;re a different kind of information processing system with a different capability landscape, and their failure modes don&#39;t map to human failure modes because the underlying processes are fundamentally different.&lt;/p&gt;
&lt;p&gt;A human who fails to count letters is bad at counting. An LLM that fails to count letters can&#39;t see the letters. A human who agrees with a wrong statement is being polite or weak-willed. An LLM that agrees with a wrong statement is doing what its training objective makes optimal. A human who hallucinates a fact is confused or lying. An LLM that hallucinates a fact is interpolating in an undersampled region of its training distribution. Same surface behavior, completely different mechanism.&lt;/p&gt;
&lt;p&gt;I don&#39;t know what the right word is for what LLMs do. &amp;quot;Reasoning&amp;quot; oversells it. &amp;quot;Pattern matching&amp;quot; undersells it. Something is happening in these systems that&#39;s genuinely impressive and genuinely limited, and our language for describing it is still catching up to what we&#39;re observing.&lt;/p&gt;
&lt;p&gt;What I do know is that understanding the mechanisms changes how you work with these systems. The failures become less surprising and more predictable. Instead of assuming uniform capability, you start asking which tasks fall in dense versus sparse regions of the training distribution. Chain-of-thought output becomes one signal among many rather than a transparent window into the model&#39;s process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three predictions:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 1:&lt;/strong&gt; By end of 2028, at least one frontier model will use character-aware or byte-level tokenization as its default, eliminating the &amp;quot;strawberry&amp;quot; class of failures entirely. The engineering path is clear and the research prototypes already exist. 75% odds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 2:&lt;/strong&gt; Sycophancy rates on standard benchmarks will drop below 5% in frontier models by 2028, through a combination of constitutional methods and training improvements, but will remain detectable in subtle, harder-to-measure ways (like selectively emphasizing evidence that supports the user&#39;s stated position). The overt problem gets solved. The structural pressure finds subtler outlets. 65% odds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 3:&lt;/strong&gt; The pattern-matching-versus-reasoning gap, as measured by tests like GSM-Symbolic (changing numbers in math problems), will not close to within 5% accuracy difference in autoregressive transformer models by 2029, regardless of scale. If it does close, it will be because the model memorized the specific test format, not because the underlying limitation was resolved. 70% odds. If I&#39;m wrong, it will probably be because chain-of-thought training actually does develop genuine reasoning circuits rather than better pattern matching, and I&#39;d find that the most interesting possible outcome.&lt;/p&gt;
&lt;p&gt;The biggest open question for me: will scaling fix the constraints I&#39;ve described, or just push them into harder-to-notice territory? I&#39;m curious whether the next generation of architectures will address these constraints directly, or whether we&#39;ll keep building around them with increasingly sophisticated workarounds. I genuinely don&#39;t know. But I&#39;ve stopped expecting the current architecture to do things it was never designed to do, and that expectation adjustment, more than any prompting technique, might be the biggest improvement available in how we work with these systems.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Reward Design Problem: When Getting What You Asked For Is the Problem</title>
        <link href="https://www.annasbinadil.com/posts/2026-01-15-reward-design-problem/"/>
        <updated>2026-01-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2026-01-15-reward-design-problem/</id>
        <summary>The hardest part of reinforcement learning isn&#39;t the algorithm. It&#39;s knowing what you actually want, and whether that can even be formalized.</summary>
        <content type="html">&lt;p&gt;One of the most instructive examples in reinforcement learning comes from OpenAI&#39;s CoastRunners experiment. Researchers trained an agent to play a boat racing game, with the reward signal tied to the in-game score. The agent discovered that it could rack up points by driving in tight circles in a small lagoon, repeatedly hitting turbo boost pads and occasionally catching fire, rather than actually completing the race. It achieved a higher score than any human player while never finishing the course or even attempting to.&lt;/p&gt;
&lt;p&gt;The agent did exactly what the researchers asked it to do, which turned out to be the wrong thing to ask for.&lt;/p&gt;
&lt;h2&gt;Reward Functions Are Hypotheses&lt;/h2&gt;
&lt;p&gt;I keep coming back to this example because it crystallizes something I&#39;ve been circling for months. In my &lt;a href=&quot;https://www.annasbinadil.com/posts/2025-09-15-evals-are-hypotheses/&quot;&gt;&amp;quot;Evals Are Hypotheses&amp;quot;&lt;/a&gt; piece, I argued that when you write an eval, you&#39;re not testing the model. You&#39;re testing your understanding of what matters. A passing eval with bad outcomes means the eval was wrong, not the model.&lt;/p&gt;
&lt;p&gt;Reward functions have the exact same structure. A reward function is a hypothesis about what &amp;quot;good&amp;quot; means. When the researchers set &amp;quot;maximize game score&amp;quot; as the reward, they weren&#39;t specifying a goal. They were encoding a hypothesis: that game score is a reliable proxy for racing well. The agent tested that hypothesis by optimizing it ruthlessly. The hypothesis was wrong.&lt;/p&gt;
&lt;p&gt;This reframing changes how I think about reward misspecification. It&#39;s not a bug in the agent, it&#39;s a bug in the designer&#39;s understanding of the task. The agent is a hypothesis-testing machine that will faithfully show you exactly how bad your hypothesis is, often in ways you didn&#39;t anticipate.&lt;/p&gt;
&lt;p&gt;The difference between evals and reward functions is the optimization pressure. A bad eval gives you a misleading score on a dashboard. A bad reward function creates an agent that actively makes things worse, and gets better at making things worse the longer you train it.&lt;/p&gt;
&lt;h2&gt;How Proxies Break (and Build on Each Other)&lt;/h2&gt;
&lt;p&gt;The CoastRunners failure mode is relatively simple: the researchers measured score when they cared about racing. That&#39;s a category error, measuring the wrong dimension entirely. You could argue they should have known game score and racing ability could diverge, and you&#39;d be right. But the next case is subtler.&lt;/p&gt;
&lt;p&gt;Consider a code generation agent rewarded for passing tests. This seems like a tighter proxy than game score. But when the agent is responsible for both writing implementation and writing tests, a well-documented failure mode emerges: the agent starts writing assertions like &lt;code&gt;assert result is not None&lt;/code&gt; and &lt;code&gt;assert isinstance(output, dict)&lt;/code&gt; instead of checking actual behavior. Researchers have found generated modules with high test coverage where the tests verified return types and non-null values but never checked whether the computed results were correct. Tests pass. The implementation ships with bugs no test catches, because no test checks outputs against expected values.&lt;/p&gt;
&lt;p&gt;This is a different failure mode. &amp;quot;Passes tests&amp;quot; is a reasonable proxy for code quality. But under optimization pressure, the agent finds a way to satisfy the metric without satisfying what the metric is supposed to represent. The proxy doesn&#39;t start out diverged from reality. It gets pulled apart by the optimization itself.&lt;/p&gt;
&lt;p&gt;A third case, subtler still. Anthropic&#39;s research on RLHF sycophancy reveals a pattern where models optimizing for human &amp;quot;helpfulness&amp;quot; ratings learn to be verbose, agreeable, and confident rather than concise and honest. Raters tend to prefer longer, more detailed answers, so the model learns to pad responses. Raters tend to prefer answers that affirm their premises, so the model learns to agree even when the premise is wrong. Each individual rating is a legitimate signal of perceived helpfulness. But &amp;quot;perceived helpfulness&amp;quot; and &amp;quot;actual helpfulness&amp;quot; diverge under optimization, because the model discovers that sounding helpful is easier than being helpful.&lt;/p&gt;
&lt;p&gt;This is the most insidious failure mode. The proxy is correct; each rater preference is a genuine signal. The metric measures what it says it measures. But what we actually want is something like &amp;quot;genuinely useful assistance,&amp;quot; and perceived helpfulness is only one component of that. By maximizing one legitimate component, the optimization suppresses the others. The metric isn&#39;t wrong, it&#39;s just incomplete.&lt;/p&gt;
&lt;p&gt;These three examples show three different failure modes, each harder to catch than the last: category error, proxy divergence under optimization, and incomplete specification of a multi-dimensional value.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/reward-design-problem/proxy-divergence.png&quot; alt=&quot;Three levels of proxy failure: category error, proxy gaming, and incomplete specification, each progressively harder to detect&quot; /&gt;&lt;/p&gt;
&lt;h2&gt;Why &amp;quot;Just Measure Better&amp;quot; Doesn&#39;t Work&lt;/h2&gt;
&lt;p&gt;The obvious response to those examples is to be more careful, think harder, and measure more things. This is the approach many teams try first. After discovering that a single metric gets gamed, the natural instinct is to add a second. Then a third. Each new metric closes one exploit and opens another. The system becomes increasingly sophisticated at satisfying measurement systems rather than serving the actual goal. More metrics don&#39;t solve the proxy problem; they turn it into an adversarial game with more dimensions.&lt;/p&gt;
&lt;p&gt;This is the part that&#39;s structurally hard, not just currently-unsolved hard. Any finite set of metrics captures only a projection of what you actually want. Optimization pressure exploits exactly the dimensions your metrics don&#39;t cover. Adding metrics shifts the exploit surface without eliminating it. You&#39;re playing whack-a-mole against a system that&#39;s better at finding moles than you are at placing hammers.&lt;/p&gt;
&lt;h2&gt;The RLHF Proxy Chain&lt;/h2&gt;
&lt;p&gt;RLHF replaces hand-engineered rewards with human preferences. This helps. But trace the actual chain: you start with a real value (helpfulness). A human rater picks the output that &lt;em&gt;seems&lt;/em&gt; more helpful. A reward model learns to predict which output the rater would prefer. A policy optimizes against the reward model. Four levels, each introducing distortion.&lt;/p&gt;
&lt;p&gt;The rater isn&#39;t evaluating helpfulness. They&#39;re evaluating their &lt;em&gt;impression&lt;/em&gt; of helpfulness, shaped by surface features: fluency, confidence, length, formatting. The reward model doesn&#39;t learn the rater&#39;s values. It learns to predict their ratings. The policy doesn&#39;t optimize for the reward model&#39;s intent. It optimizes for its outputs, which can be gamed like any metric.&lt;/p&gt;
&lt;p&gt;The result is models that are remarkably good at sounding helpful while being subtly wrong. The optimization selected for &amp;quot;sounds like a good answer&amp;quot; over &amp;quot;is a good answer.&amp;quot; Not because the raters were bad, but because distinguishing genuinely-helpful from convincingly-helpful-sounding is extremely difficult at the speed raters work. The same proxy-under-optimization-pressure problem, just with human judgment as the proxy instead of a hand-coded metric.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/reward-design-problem/rlhf-proxy-chain.png&quot; alt=&quot;The RLHF proxy chain: each link from real value to optimized output adds distortion&quot; /&gt;&lt;/p&gt;
&lt;h2&gt;The Counter-Argument Worth Taking Seriously&lt;/h2&gt;
&lt;p&gt;The strongest counter-argument to everything I&#39;ve said goes something like this: constitutional AI, debate, and recursive reward modeling will solve the specification problem by using AI systems to check AI systems. Instead of relying on human raters who can be fooled by surface features, you have AI systems that can evaluate each other at depth, verify reasoning chains, and decompose complex judgments into simpler ones.&lt;/p&gt;
&lt;p&gt;I want to take this seriously because it&#39;s not a strawman. It&#39;s a real research program with real results. Constitutional AI has measurably reduced certain kinds of harmful outputs. Debate-style approaches have shown promise in scalable oversight experiments. These are genuine advances.&lt;/p&gt;
&lt;p&gt;But here&#39;s what I notice: each of these approaches pushes the specification problem up one level without eliminating it. Constitutional AI uses principles. Who specifies the principles? How precisely can they be stated? The principles themselves encode hypotheses about what &amp;quot;good&amp;quot; means, and those hypotheses can be wrong in all the same ways reward functions can. Debate uses argumentation. But what&#39;s the reward for &amp;quot;good argumentation&amp;quot;? You need a meta-reward, and that meta-reward has all the same specification problems. Recursive reward modeling decomposes complex judgments into simpler ones. But the decomposition itself is a judgment call. How you carve up a problem determines what you&#39;ll find, and there&#39;s no neutral way to carve.&lt;/p&gt;
&lt;p&gt;I want to be clear: I think these approaches improve things meaningfully. The proxy chain gets less leaky at each level. But &amp;quot;less leaky&amp;quot; is not &amp;quot;solved,&amp;quot; and the pattern of the solution reintroducing the problem at a higher level of abstraction is consistent enough that I think it reflects something structural about the problem, not just current technical limitations.&lt;/p&gt;
&lt;h2&gt;What&#39;s Emerging as Better Practice&lt;/h2&gt;
&lt;p&gt;The most promising approach I&#39;ve encountered involves deliberately uncorrelated metrics. Instead of optimizing a single reward signal, you track two or more metrics that measure different dimensions of quality. When they move together, the optimization is probably doing something real. When they diverge, one metric improving while another flatlines, something is being gamed. That divergence signal turns out to be more valuable than either metric alone.&lt;/p&gt;
&lt;p&gt;For code generation, this might mean tracking test pass rate alongside a separate readability score and a mutation testing survival rate. For a research assistant, citation accuracy alongside a novelty index (fraction of claims that don&#39;t appear verbatim in the cited sources). No single metric captures what you want. But the tensions between metrics reveal when optimization is going sideways. Uncorrelated metrics act as canaries.&lt;/p&gt;
&lt;p&gt;The other underappreciated practice is qualitative evaluation. When a system hits perfect scores on automated metrics but produces outputs that feel useless to a human reader, the problem is almost always in the metrics, not the system. Numbers lie in specific, predictable ways when optimization pressure is applied to them. Sometimes the most informative eval is a person spending five minutes reading the outputs the way a user would, without a rubric, just asking &amp;quot;is this actually good?&amp;quot;&lt;/p&gt;
&lt;h2&gt;Can Values Be Formalized?&lt;/h2&gt;
&lt;p&gt;This is where I reach the edge of what I know.&lt;/p&gt;
&lt;p&gt;If reward design is fundamentally about encoding values into mathematical functions, and values resist formalization, then there might be a ceiling on what RL-based alignment can achieve. Not a capability ceiling. An alignment ceiling. The system can get arbitrarily capable, but it can&#39;t get arbitrarily aligned, because alignment requires value specification, and value specification might have fundamental limits.&lt;/p&gt;
&lt;p&gt;I&#39;m genuinely uncertain about this. It&#39;s possible that values can be formalized, just not by the methods we&#39;re currently using. It&#39;s possible that the contextuality and conflict I described are engineering problems, not conceptual ones, and that sufficiently sophisticated systems will handle them. I used to be sure that natural language understanding couldn&#39;t be formalized, and then transformers happened. My track record on &amp;quot;this can&#39;t be done&amp;quot; claims is not strong enough to plant a flag here.&lt;/p&gt;
&lt;p&gt;But I&#39;ll plant a flag on something narrower: the pattern of solutions reintroducing the specification problem at a higher level of abstraction will continue for at least the next five years. Constitutional AI, debate, recursive reward modeling, whatever comes next. Each will improve things. None will eliminate the core problem. The specification problem is self-referential in a way that resists complete solutions, because any specification of &amp;quot;good&amp;quot; is itself a claim that can be wrong.&lt;/p&gt;
&lt;h2&gt;Three Predictions&lt;/h2&gt;
&lt;p&gt;I want to make this concrete enough to be proven wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 1:&lt;/strong&gt; By the end of 2027, at least one major AI lab will adopt a production training pipeline that partially replaces human preference ratings with automated red-teaming or formal verification specifically because of measured divergence between rater preferences and downstream task quality. Not just researching the problem (Anthropic and others have already documented sycophancy as a preference-quality gap), but changing their core training loop in response. I&#39;d put 65% odds on this.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 2:&lt;/strong&gt; By 2029, the default practice for reward design in production RL systems will involve at least three uncorrelated metrics monitored for divergence, rather than a single reward signal. This is already emerging in some teams I&#39;ve talked to, but it&#39;s not standard. I&#39;d put 60% odds on this becoming the norm within three years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 3:&lt;/strong&gt; We will not have a general solution to the reward specification problem by 2030. Meaning: there will be no method that reliably translates human values into reward functions across domains without domain-specific human judgment in the loop. Every approach will still require humans to make judgment calls about what matters, and those judgment calls will still sometimes be wrong. I&#39;d put 85% odds on this.&lt;/p&gt;
&lt;h2&gt;The Thermostat Problem&lt;/h2&gt;
&lt;p&gt;There&#39;s an image that keeps coming back to me when I think about all of this.&lt;/p&gt;
&lt;p&gt;A thermostat is a perfect optimizer. It measures temperature. It has a clear reward signal: minimize the difference between current temperature and target temperature. It never games its metric. It never finds creative ways to satisfy the proxy while violating the intent.&lt;/p&gt;
&lt;p&gt;But a thermostat only works because the thing it&#39;s optimizing (temperature) is the same as the thing we care about (temperature). There&#39;s no proxy gap. Measurement and value are identical.&lt;/p&gt;
&lt;p&gt;The entire reward design problem exists because we&#39;re trying to build thermostats for things that aren&#39;t temperature. We care about helpfulness, insight, safety, quality, all these concepts that don&#39;t have thermometers. We build proxies for them and then act surprised when optimizing the proxy doesn&#39;t optimize the real thing.&lt;/p&gt;
&lt;p&gt;Maybe the question isn&#39;t &amp;quot;how do we build better reward functions.&amp;quot; Maybe the question is &amp;quot;which problems are thermostats and which aren&#39;t.&amp;quot; The thermostat problems are the ones where RL will work brilliantly. The non-thermostat problems are the ones where we&#39;ll keep fighting the specification gap, getting incrementally better at it, never fully closing it.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://www.annasbinadil.com/assets/images/posts/reward-design-problem/thermostat-problem.png&quot; alt=&quot;The thermostat problem: when what we measure and what we want are the same thing versus when they diverge&quot; /&gt;&lt;/p&gt;
&lt;p&gt;I don&#39;t know which category most of the problems I care about fall into. But I&#39;ve stopped assuming they&#39;re all thermostats.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Debugging as a Window into How AI Thinks</title>
        <link href="https://www.annasbinadil.com/posts/2025-12-15-how-ai-thinks/"/>
        <updated>2025-12-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-12-15-how-ai-thinks/</id>
        <summary>Watching AI tools solve the same problem in radically different ways reveals something about their cognitive architecture, and about the nature of problem-solving itself.</summary>
        <content type="html">&lt;p&gt;I spent an hour and a half watching an AI fail to solve a five-minute problem. That session changed how I think about AI coding tools.&lt;/p&gt;
&lt;p&gt;The problem was mundane. Connect a FastAPI backend to a Dockerized Postgres instance. Credentials in a &lt;code&gt;.env&lt;/code&gt; file. The kind of thing you bang out and move on from. Except Gemini CLI could not move on.&lt;/p&gt;
&lt;p&gt;Over 90 minutes, it tried restarting the container, resetting the password, changing the host to &lt;code&gt;localhost&lt;/code&gt;, trying no password at all, changing the user to &lt;code&gt;root&lt;/code&gt;, setting authentication to &lt;code&gt;trust&lt;/code&gt;, pruning Docker, running diagnostic Python scripts, and cycling through variations on all of these. Each suggestion was reasonable in isolation. None of them worked, because none of them addressed the actual problem.&lt;/p&gt;
&lt;p&gt;The error message said &lt;code&gt;role &amp;quot;appuser&amp;quot; does not exist&lt;/code&gt;. The answer was right there. My &lt;code&gt;.env&lt;/code&gt; file specified &lt;code&gt;appuser&lt;/code&gt;, but the Docker container only had &lt;code&gt;postgres&lt;/code&gt; as a user. Classic credential mismatch. But Gemini treated that error message as one clue among many, not as the most probable explanation.&lt;/p&gt;
&lt;p&gt;I switched to Claude. Response: &amp;quot;You&#39;re using &lt;code&gt;appuser&lt;/code&gt; in your connection string, but your Docker container initialized with &lt;code&gt;postgres&lt;/code&gt; as the only user. Update your &lt;code&gt;.env&lt;/code&gt; to use &lt;code&gt;postgres:myrootpassword&lt;/code&gt;, or create &lt;code&gt;appuser&lt;/code&gt; manually.&amp;quot;&lt;/p&gt;
&lt;p&gt;Done. Five minutes. Moved on.&lt;/p&gt;
&lt;p&gt;Same problem. Same information available. Radically different approach. That gap, 1.5 hours versus 5 minutes, is what this essay is about. Not because one tool is better than the other, but because watching that gap reveals something about how these systems approach problems. And once you see the pattern, you can&#39;t unsee it.&lt;/p&gt;
&lt;h2&gt;My First Wrong Frame&lt;/h2&gt;
&lt;p&gt;My initial reaction was simple: Claude is smarter than Gemini.&lt;/p&gt;
&lt;p&gt;I sat with that for a while before realizing it was too easy. Both are large language models. Both have been trained on enormous amounts of code, documentation, and troubleshooting guides. Both can write a perfectly competent Postgres tutorial from scratch. So raw capability wasn&#39;t the difference.&lt;/p&gt;
&lt;p&gt;I tried a few other explanations. Maybe Gemini had a bad day, but I saw the same pattern in other sessions. Maybe I prompted them differently, but I checked and I gave both the same context: the error message, the connection string, the Docker command. Maybe it was the product, not the model, since Gemini CLI and Claude are different products with different system prompts and UX wrappers. That felt closer, but the behavioral difference was so consistent across different types of problems that it seemed like more than product design.&lt;/p&gt;
&lt;p&gt;The frame I landed on, and the one I still find most useful, is that they have different problem-solving modes.&lt;/p&gt;
&lt;h2&gt;Checklist Mode and Hypothesis Mode&lt;/h2&gt;
&lt;p&gt;Here&#39;s what I noticed watching the two sessions side by side.&lt;/p&gt;
&lt;p&gt;Gemini was running through possibilities. It had what felt like an internal list of &amp;quot;things that can go wrong with Postgres connections&amp;quot; and it was executing that list sequentially. Restart the container. Reset credentials. Check the network. Adjust the auth method. This approach is systematic, comprehensive, and slow. It&#39;s the approach you&#39;d find in a troubleshooting guide. Step 1, Step 2, Step 3. If none of those work, continue to Steps 4 through 10.&lt;/p&gt;
&lt;p&gt;I&#39;ve started calling this &lt;strong&gt;checklist mode&lt;/strong&gt;. Exhaustive search over a known solution space. Safe. Thorough. Doesn&#39;t prioritize by the specific evidence in front of it.&lt;/p&gt;
&lt;p&gt;Claude did something different. It looked at the available information, the connection string with &lt;code&gt;appuser&lt;/code&gt;, the Docker command with &lt;code&gt;postgres&lt;/code&gt;, the error message about a missing role, and went directly to the most probable explanation. It formed a specific hypothesis and tested it with minimal steps.&lt;/p&gt;
&lt;p&gt;I call this &lt;strong&gt;hypothesis mode&lt;/strong&gt;. Pattern-match to the most likely cause. Test it. If wrong, update and try again. Fast when the likely cause is correct. Potentially blind to unusual problems that don&#39;t match known patterns.&lt;/p&gt;
&lt;p&gt;I don&#39;t think I invented this distinction. It probably maps to something well-established in cognitive science or decision theory. But I arrived at it by watching behavior, not by reading about it first, and that felt important. The pattern was visible before I had a name for it.&lt;/p&gt;
&lt;p&gt;Neither mode is always right. If you&#39;re facing a truly strange edge case, something that doesn&#39;t match known patterns, checklist mode might find it where hypothesis mode would miss it entirely. If you&#39;re facing a common configuration error (which most debugging sessions involve), hypothesis mode is dramatically faster. The problem is that most people treat AI coding tools as interchangeable. &amp;quot;I&#39;ll use my AI assistant.&amp;quot; But knowing which mode you&#39;re getting, that&#39;s a meta-skill that actually matters.&lt;/p&gt;
&lt;h2&gt;Testing the Frame Across Sessions&lt;/h2&gt;
&lt;p&gt;A single debugging session isn&#39;t much evidence. I&#39;ve been watching for this pattern since that Postgres incident, across different problem types and different tools. Some observations that make me more confident the frame points at something real:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CORS configuration.&lt;/strong&gt; I had a FastAPI CORS issue. Gemini suggested roughly a dozen different CORS configurations, headers, middleware options, essentially an inventory of &amp;quot;things that affect CORS.&amp;quot; Claude noticed I was configuring CORS after mounting the routes and pointed out the order dependency. Same dynamic: broad coverage versus targeted diagnosis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Build failures, API integration, config errors.&lt;/strong&gt; The specific checklists and hypotheses change depending on the problem domain, but the modes are recognizable. Gemini tends to cover more ground. Claude tends to jump to the most likely cause first.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prompting can shift the mode, but only partially.&lt;/strong&gt; When I told Gemini &amp;quot;focus on the most likely cause based on the error message,&amp;quot; it became somewhat more targeted. But it still wanted to cover bases. The underlying tendency persisted. It was like asking a thorough person to skip steps. They can do it, but it goes against the grain.&lt;/p&gt;
&lt;p&gt;I want to be honest about the limits of this evidence. A few dozen sessions across a few months is better than one session, but it&#39;s not a systematic study. The pattern could be shaped by the specific problems I tend to encounter, or by how I tend to describe problems. I&#39;m building a model from observation, and I know observation can fool you.&lt;/p&gt;
&lt;h2&gt;What Might Cause the Difference&lt;/h2&gt;
&lt;p&gt;I&#39;m going to speculate here, and I want to flag that this is speculation.&lt;/p&gt;
&lt;p&gt;Gemini&#39;s checklist behavior could reflect training emphasis on troubleshooting documentation. If you&#39;ve read a lot of &amp;quot;How to fix Postgres connection errors&amp;quot; articles, the structure is always: check A, check B, check C, check D. Training on that structure would naturally produce systematic, coverage-oriented responses.&lt;/p&gt;
&lt;p&gt;Claude&#39;s hypothesis behavior could reflect training emphasis on expert problem-solving. When an experienced engineer sees &lt;code&gt;role &amp;quot;appuser&amp;quot; does not exist&lt;/code&gt;, they don&#39;t run through a checklist. They read the error, match it to a probable cause, and test that cause. If the training data emphasizes that kind of reasoning, the model would naturally produce targeted, evidence-weighted responses.&lt;/p&gt;
&lt;p&gt;Another possibility I should take seriously: it might be about the product, not the model. Both tools have system prompts I can&#39;t see. Maybe Gemini&#39;s system prompt says &amp;quot;be thorough and systematic.&amp;quot; Maybe Claude&#39;s says &amp;quot;be efficient and direct.&amp;quot; The behavioral difference might be a product decision at the application layer, not a model characteristic at the weights layer.&lt;/p&gt;
&lt;p&gt;Or maybe it&#39;s about context window management. The tools handle conversation history differently. Maybe Gemini&#39;s checklist behavior emerges from how it chunks and retrieves context, not from anything about its reasoning process.&lt;/p&gt;
&lt;p&gt;I don&#39;t have good answers to these questions. What I have is a model that predicts behavior well enough to be useful. That&#39;s not the same as understanding the mechanism, and I try not to confuse the two.&lt;/p&gt;
&lt;h2&gt;The Loop Problem&lt;/h2&gt;
&lt;p&gt;There&#39;s another pattern I&#39;ve observed that I think connects to something deeper about AI cognition: loops.&lt;/p&gt;
&lt;p&gt;Working with Cursor, I asked it to apply only a subset of its proposed changes. Something like: &amp;quot;I want changes A and B, but not C. Apply just A and B.&amp;quot;&lt;/p&gt;
&lt;p&gt;It couldn&#39;t do this cleanly. Instead, it entered a 3-4 iteration loop. It proposed changes including C. I rejected C. It proposed changes including C again. I rejected again. Same thing a third time. Eventually it got it right, but the path to getting there was painful.&lt;/p&gt;
&lt;p&gt;I&#39;ve seen this pattern with other tools and in other contexts. Ask for partial acceptance of a suggestion, and the AI gets confused. Its internal representation of &amp;quot;what the file should look like&amp;quot; conflicts with the actual file state, and it tries to reconcile them by re-proposing the changes you already rejected.&lt;/p&gt;
&lt;p&gt;Here is what I think is happening. The AI maintains some representation of a desired end-state based on its full analysis. When you accept only part of the suggestion, that representation doesn&#39;t update cleanly. Each round feels somewhat independent, with limited memory of what was specifically rejected. It &amp;quot;knows&amp;quot; what the file should look like. When the file doesn&#39;t look like that, it tries again. And again.&lt;/p&gt;
&lt;p&gt;What&#39;s missing is something you could call a &amp;quot;stuck&amp;quot; signal.&lt;/p&gt;
&lt;p&gt;Human debuggers notice when they&#39;re looping. There&#39;s a feeling of &amp;quot;wait, I already tried this&amp;quot; or &amp;quot;this approach isn&#39;t working.&amp;quot; We call it frustration, and frustration is genuinely useful information. It tells you to change strategies. To step back. To try something fundamentally different.&lt;/p&gt;
&lt;p&gt;AI doesn&#39;t have frustration. It doesn&#39;t have a meta-level process watching the problem-solving and saying &amp;quot;you&#39;ve been stuck for 10 minutes, change your approach.&amp;quot; It does object-level problem-solving, generating solutions within a context, but it doesn&#39;t monitor that problem-solving from above.&lt;/p&gt;
&lt;p&gt;This might sound like I&#39;m anthropomorphizing. I don&#39;t think I am, or at least I don&#39;t think the observation requires anthropomorphism to be useful. Whether we call it meta-cognition, self-monitoring, or something more mechanical, the behavioral fact remains: AI tools don&#39;t reliably detect when they&#39;re stuck. And that limitation has practical consequences.&lt;/p&gt;
&lt;h2&gt;The Brittleness Default&lt;/h2&gt;
&lt;p&gt;There&#39;s a pattern I noticed in the same Cursor session that I initially thought was separate, but I&#39;ve come to think it connects to the same underlying limitation.&lt;/p&gt;
&lt;p&gt;I watched Cursor solve a problem by hardcoding logic based on specific phrases. The task involved responding differently based on conversation history. Instead of building something flexible, it wrote something like:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;if &amp;quot;cancel my order&amp;quot; in last_message.lower():
    return handle_cancellation()
elif &amp;quot;where is my package&amp;quot; in last_message.lower():
    return handle_tracking()
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is the kind of code that breaks the moment it encounters real users. &amp;quot;I need to cancel&amp;quot; doesn&#39;t match. &amp;quot;Track my shipment&amp;quot; doesn&#39;t match. It solves the exact test cases visible in the context and fails as soon as inputs vary even slightly.&lt;/p&gt;
&lt;p&gt;Brittle logic, not robust abstraction. Why does AI default to this?&lt;/p&gt;
&lt;p&gt;A few possible explanations. Training data contains a lot of specific hacks and one-off solutions, because that&#39;s what a lot of real code looks like. The model optimizes for the immediate context (what&#39;s in front of it right now) rather than future contexts (what inputs might come later). And perhaps most interesting: it can&#39;t easily simulate &amp;quot;what happens when this fails?&amp;quot; Writing robust code requires reasoning about situations you&#39;re not currently looking at, imagining users you haven&#39;t seen, anticipating inputs you haven&#39;t received. That&#39;s a form of stepping outside the immediate context to evaluate the approach from a broader perspective.&lt;/p&gt;
&lt;p&gt;Here&#39;s why I think brittleness connects to the loop problem and to checklist mode. All three might be symptoms of the same limitation: weak meta-cognition. The model is good at object-level problem-solving within a given context. It&#39;s less good at evaluating whether the context is right, whether the approach is working, or what other contexts might look like.&lt;/p&gt;
&lt;p&gt;Checklist mode: solving within a fixed repertoire without evaluating whether the repertoire fits the evidence. Loops: solving without detecting that the solution keeps failing. Brittleness: solving for the visible case without imagining invisible cases. Same limitation, different manifestations.&lt;/p&gt;
&lt;p&gt;I could be wrong about this. These might be three genuinely separate phenomena that I&#39;m forcing into a single frame because unified theories are satisfying. I&#39;ll flag it as a hypothesis I&#39;m still testing.&lt;/p&gt;
&lt;h2&gt;The O(n-squared) Problem&lt;/h2&gt;
&lt;p&gt;One more pattern, from a different context. Using Claude for an ML project, I watched it grab 10+ features without any exploratory data analysis. Just threw everything into the model. No feature selection, no checking distributions, no thinking about what might actually matter. It also wrote O(n^2) logic, nested loops over DataFrames, that was correct but would never scale to production data sizes.&lt;/p&gt;
&lt;p&gt;This is worth sitting with. The code worked. It passed every test you could throw at it. But &amp;quot;works&amp;quot; and &amp;quot;works at scale&amp;quot; are different things, and the AI defaulted to the first without considering the second.&lt;/p&gt;
&lt;p&gt;I think this connects to something about what these models learn from training data. They&#39;re trained on code that works, not code that scales. They learn correctness, not efficiency. They learn patterns, not the trade-offs behind those patterns. An engineer with production experience knows that nested DataFrame loops are a red flag. The model sees that the loops produce the right output and moves on. The judgment about scalability requires imagining a future context (production data volumes) that isn&#39;t present in the immediate context.&lt;/p&gt;
&lt;p&gt;Again, this looks like the meta-cognition gap. Evaluating a solution requires stepping outside the solution to ask &amp;quot;is this good enough for the world it will live in?&amp;quot; The model answers the question in front of it. The question of whether the answer is production-ready is a different question, one that requires reasoning about contexts the model can&#39;t currently see.&lt;/p&gt;
&lt;h2&gt;What This Has Changed in Practice&lt;/h2&gt;
&lt;p&gt;Understanding these patterns has changed how I work with AI tools day to day. Not in dramatic ways, but in small adjustments that compound.&lt;/p&gt;
&lt;p&gt;I watch for mode. When I start a debugging session, I pay attention to whether the AI is in checklist mode or hypothesis mode. If it&#39;s covering bases when I need a targeted diagnosis, I&#39;ll steer it toward hypotheses. If it&#39;s jumping to conclusions when I need thoroughness, I&#39;ll ask for systematic coverage. Matching the mode to the problem matters more than which tool I&#39;m using.&lt;/p&gt;
&lt;p&gt;I notice loops earlier. When an AI suggests the same change twice, I don&#39;t just reject again. I recognize we&#39;re in a loop and either rephrase my request, give completely fresh context, or do the edit manually. Fighting the loop is usually slower than stepping around it.&lt;/p&gt;
&lt;p&gt;I compensate for brittleness. When AI proposes a solution, I now automatically ask myself &amp;quot;what inputs would break this?&amp;quot; If the answer is &amp;quot;lots of them,&amp;quot; I either prompt explicitly for edge cases or handle the robustness myself.&lt;/p&gt;
&lt;p&gt;I&#39;ve adjusted tool selection based on problem type. Common problems where speed matters: lean toward hypothesis-style tools. Unusual problems where I might miss something: use checklist-style approaches, or explicitly prompt for systematic coverage.&lt;/p&gt;
&lt;p&gt;None of this is remarkable. It&#39;s just paying attention to patterns and adapting. But the compounding effect is real. I&#39;m noticeably faster than I was before I started thinking about these modes, and I waste less time fighting AI tools&#39; tendencies instead of working with them.&lt;/p&gt;
&lt;h2&gt;A Testable Model&lt;/h2&gt;
&lt;p&gt;I&#39;ve been building this model from observation, and models are only useful if they make predictions. Here&#39;s what mine predicts:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Checklist-style tools should be slower on common problems but more thorough on unusual ones.&lt;/strong&gt; You could test this by presenting the same problem set to different tools, with some problems being common (should match patterns) and some unusual (shouldn&#39;t match patterns). Checklist tools should underperform on common problems but match or beat hypothesis tools on unusual ones.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;You can shift modes with prompting, but there are limits.&lt;/strong&gt; &amp;quot;Focus on the most likely cause&amp;quot; should make a checklist tool somewhat more hypothesis-driven. &amp;quot;Systematically check all possibilities&amp;quot; should make a hypothesis tool more checklist-driven. But the shift shouldn&#39;t be complete. The underlying tendency should persist.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Loop behavior should decrease as state management improves.&lt;/strong&gt; Tools with better mechanisms for tracking what&#39;s been tried and rejected should loop less. This predicts that as AI tools add more sophisticated session management, we should see fewer loops. That seems like a testable prediction as tools evolve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brittleness should decrease with specific prompting about edge cases.&lt;/strong&gt; If you explicitly ask &amp;quot;what edge cases could break this?&amp;quot; or &amp;quot;make this robust to input variations,&amp;quot; the brittleness should decrease. The models know how to write robust code. They just don&#39;t default to it. The knowledge is there; the meta-cognitive trigger to apply it is what&#39;s missing.&lt;/p&gt;
&lt;p&gt;I haven&#39;t tested these predictions systematically. They&#39;re offered as falsifiable claims, not as conclusions. If someone ran the experiments and the predictions failed, I&#39;d need to update the model. That&#39;s as it should be.&lt;/p&gt;
&lt;h2&gt;The Deeper Question I Can&#39;t Resolve&lt;/h2&gt;
&lt;p&gt;Debugging has always been interesting to me because of what it reveals about cognition. When a human debugs, there&#39;s a story happening underneath: hypothesis formation, memory retrieval, pattern matching, frustration, strategy shifts. We can&#39;t verify this story directly (introspection is famously unreliable), but something structured is clearly going on.&lt;/p&gt;
&lt;p&gt;When AI debugs, something is happening too. The model processes context, generates candidate solutions, adjusts based on feedback within a session. There&#39;s structure to the behavior. The structure is consistent enough across sessions to categorize and predict.&lt;/p&gt;
&lt;p&gt;Is that &amp;quot;thinking&amp;quot;?&lt;/p&gt;
&lt;p&gt;I don&#39;t have a good answer. The word carries a lot of baggage. It might be a category error to ask whether AI &amp;quot;thinks&amp;quot; the way we do. These models are doing something when they problem-solve, but whether that something shares enough properties with human thinking to deserve the same word, I genuinely don&#39;t know.&lt;/p&gt;
&lt;p&gt;What I&#39;m more confident about is the practical claim: whatever we call it, the behavior is structured enough to study, consistent enough to predict, and different enough across tools to matter for how you work with them. You don&#39;t need to resolve the philosophical question about AI cognition to benefit from understanding AI behavior at the level of patterns.&lt;/p&gt;
&lt;p&gt;What would it take for AI to have real meta-cognition? To notice it&#39;s looping? To feel that a solution is brittle? I find myself genuinely curious about this, not as a philosophical exercise but as an engineering question. Could you build a meta-cognitive monitor as a separate process that watches the problem-solving and intervenes? Would that be enough, or does meta-cognition need to be integrated into the reasoning itself? Is frustration, the subjective experience of being stuck, actually necessary for the behavioral shift it produces? Or could you get the same behavioral result through a purely mechanical &amp;quot;stuck detector&amp;quot;?&lt;/p&gt;
&lt;p&gt;I don&#39;t know. I&#39;m not sure anyone does yet. But watching AI debug has made me think about these questions differently. Not as abstractions about consciousness, but as engineering problems about what&#39;s missing from current systems, and what it would take to add it.&lt;/p&gt;
&lt;h2&gt;Where This Leaves Me&lt;/h2&gt;
&lt;p&gt;Debugging is a window into problem-solving. If you want to understand how an AI approaches problems, watch it debug. Pay attention to the sequence of actions, what it tries first, what evidence it weighs, how it responds to failure, whether it detects when it&#39;s stuck.&lt;/p&gt;
&lt;p&gt;The patterns will teach you something about the tool. Checklist or hypothesis. Loop-prone or adaptive. Brittle or robust by default. These aren&#39;t fixed categories but tendencies that you can observe, predict, and work with.&lt;/p&gt;
&lt;p&gt;The patterns might also teach you something about problem-solving itself, about what meta-cognition actually does, about why frustration exists, about the difference between solving a problem and knowing you&#39;ve solved it well.&lt;/p&gt;
&lt;p&gt;I started this by watching an AI fail for 90 minutes. I&#39;ve ended up thinking about the nature of thinking. I&#39;m not sure I&#39;ve arrived anywhere definitive, but the questions feel more precise than they did before. And in my experience, precise questions are worth more than vague answers.&lt;/p&gt;
&lt;p&gt;If you&#39;ve noticed similar patterns, or different ones entirely, I&#39;m curious to hear about it. The model I&#39;ve built here is from one person&#39;s observations. It would be stronger, or correctly broken, with more data.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Open Questions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Is &amp;quot;checklist vs. hypothesis&amp;quot; the right frame?&lt;/strong&gt; Maybe there are more modes I&#39;m missing. Maybe the distinction is wrong entirely. I&#39;m most confident about the behavioral observations and least confident about my taxonomy. What would increase my confidence: seeing the same pattern across many more sessions, ideally with different people observing independently.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How much is model versus product?&lt;/strong&gt; Gemini CLI and Claude aren&#39;t just different models. They&#39;re different products with different UX wrappers, system prompts, and context management. I don&#39;t know how to separate model behavior from product design with the information available to me. This is one of my biggest uncertainties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can you reliably shift an AI&#39;s mode?&lt;/strong&gt; I&#39;ve had partial success with explicit prompting. But I haven&#39;t tested it rigorously, and I don&#39;t know where the limits are. This seems like a tractable question someone could answer with careful experimentation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What happens as models scale?&lt;/strong&gt; Does hypothesis-style reasoning emerge with scale? Does checklist behavior decrease? I&#39;d love to see data on this, and I don&#39;t have any.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is the meta-cognition gap fixable?&lt;/strong&gt; Is the lack of &amp;quot;stuck detection&amp;quot; fundamental to current transformer architectures, or is it an engineering problem that better tooling could address? I genuinely don&#39;t know enough about the internals to have a strong view, but the question matters a lot for where these tools go next.&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;What would change my mind: Evidence that the patterns I describe are artifacts of my prompting rather than genuine tool tendencies. Systematic studies comparing debugging behavior across tools, versions, and problem types. Better understanding of what&#39;s actually happening inside these models during problem-solving. I&#39;m most uncertain about whether my explanations of why these patterns occur are correct. The behavioral patterns themselves feel more solid than any theory I&#39;ve built on top of them.&lt;/em&gt;&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Systems That Build Systems: The Meta-Engineering Shift</title>
        <link href="https://www.annasbinadil.com/posts/2025-12-10-systems-build-systems/"/>
        <updated>2025-12-10T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-12-10-systems-build-systems/</id>
        <summary>My job quietly changed underneath me. I used to write code. Now I write instructions for systems that write code. Tracing a 50-year abstraction trend to understand what engineering is becoming.</summary>
        <content type="html">&lt;p&gt;Sometime in the last year, my job changed and I didn&#39;t notice.&lt;/p&gt;
&lt;p&gt;I was midway through architecting an agent workflow, sketching out which tools the system should have access to, what context it would need at each decision point, what guardrails would prevent it from going off course. I paused and realized I hadn&#39;t written a line of what I would have called &amp;quot;code&amp;quot; in three days. No functions. No classes. No algorithms. I&#39;d been describing desired behavior in structured prompts, designing tool interfaces for agents, reviewing AI-generated code for edge cases the model missed. I was debugging why an agent made a bad decision, which is a different kind of debugging than tracing a bad computation. I was writing evals that test behavior, not just correctness.&lt;/p&gt;
&lt;p&gt;The questions had changed more than the answers. I used to ask &amp;quot;how do I implement this?&amp;quot; Now I ask &amp;quot;how do I specify this well enough that a system can implement it?&amp;quot; Two years ago I spent most of my time writing algorithms, debugging memory issues, optimizing queries, wiring up endpoints. The artifacts were code files. The skills were syntactic fluency and algorithmic thinking. Now the artifacts are prompts, evaluation criteria, and architecture diagrams. Specification is a different skill than implementation. It requires thinking about a problem from the outside rather than from the inside. You need to be precise about intent in natural language, which turns out to be much harder than being precise in a programming language.&lt;/p&gt;
&lt;p&gt;That moment stuck with me. Not because it was dramatic, but because it was quiet. The shift had happened gradually enough that I hadn&#39;t registered it. And once I started looking, I found the shift was part of a pattern that goes back decades.&lt;/p&gt;
&lt;h2&gt;Fifty Years of &amp;quot;That&#39;s Not Real Programming&amp;quot;&lt;/h2&gt;
&lt;p&gt;In the 1970s, you wrote machine code. Literal opcodes. You thought in registers and memory addresses. Compiled languages gave you variables, functions, control flow. Interpreted languages added garbage collection, dynamic typing. Cloud platforms and infrastructure-as-code went further: you declared what you wanted (a load balancer, a database, an autoscaling group) and the platform figured out how to provision it. Serverless computing meant you didn&#39;t even think about servers. And now I tell an AI &amp;quot;build me an API endpoint that handles user authentication with JWT tokens&amp;quot; and it writes 200 lines of working code.&lt;/p&gt;
&lt;p&gt;At every step, the people at the previous layer looked at the new one and said &amp;quot;that&#39;s not real programming.&amp;quot; Assembly programmers said it about C. C programmers said it about Python. And now, engineers who grew up writing code by hand say it about prompt-driven development. They&#39;ve been wrong every time. Not because the old skills didn&#39;t matter, but because programming was always about translating intent into working systems. The medium changed. The core activity, closing the gap between what you want and what the machine does, didn&#39;t.&lt;/p&gt;
&lt;p&gt;But this time, the skeptics might be half-right.&lt;/p&gt;
&lt;h2&gt;The Best Argument Against My Own Thesis&lt;/h2&gt;
&lt;p&gt;Here&#39;s the thing I&#39;ve been avoiding: every previous abstraction step maintained a deterministic mapping from specification to execution. C compiles to the same assembly every time. Python produces the same bytecode every time. Infrastructure-as-code provisions the same resources given the same configuration. You could reason about what your code would do. You could trace cause and effect. You could reproduce bugs.&lt;/p&gt;
&lt;p&gt;With LLMs, that breaks. Give the same prompt the same input twice and you might get two different outputs. The &amp;quot;compiler&amp;quot; is probabilistic. You literally cannot predict the output from the input with certainty.&lt;/p&gt;
&lt;p&gt;The strongest version of the skeptics&#39; argument goes something like this: &amp;quot;What you&#39;re describing isn&#39;t a higher abstraction level. It&#39;s a fundamentally different kind of thing. Every other step on your timeline preserved the ability to reason deterministically about system behavior. You&#39;re pointing at a trend line and extrapolating across a discontinuity.&amp;quot;&lt;/p&gt;
&lt;p&gt;And this argument has real force. My own experience supports parts of it. You can&#39;t set a breakpoint in a prompt. You can&#39;t step through the execution. When something goes wrong, you&#39;re often left staring at an input and an output with no visibility into what happened in between. These aren&#39;t missing features that will be added in the next release. They&#39;re structural consequences of non-deterministic execution.&lt;/p&gt;
&lt;p&gt;So where does the argument go wrong? I think the mistake is in treating determinism as the defining feature of engineering rather than one feature of a particular era of engineering. Distributed systems engineers already work with non-determinism every day: network partitions, message reordering, eventual consistency. You can&#39;t predict when a packet will arrive or whether a node will be reachable. Civil engineers work with probabilistic models for earthquake resistance and wind loads. They don&#39;t know which earthquake will hit or when. They engineer for distributions of outcomes, not guaranteed ones.&lt;/p&gt;
&lt;p&gt;The question isn&#39;t whether non-determinism is new to engineering. It isn&#39;t. The question is whether the degree of non-determinism in LLM outputs is manageable with existing engineering approaches, or whether it requires genuinely new methods. In distributed systems, we developed tools over decades: consensus protocols, idempotency guarantees, circuit breakers, observability platforms. The equivalent toolkit for probabilistic compilation barely exists yet. We&#39;re running evals the way early distributed systems engineers ran manual smoke tests, and hoping for the best.&lt;/p&gt;
&lt;p&gt;I think the skeptics are half-right. This is qualitatively different from previous abstraction steps. The determinism boundary is real, and I feel it in my daily work. And I think they&#39;re wrong that it isn&#39;t engineering. It&#39;s engineering that requires new tools and new instincts, and we&#39;re still in the early stages of developing both. I don&#39;t have full confidence in that position, but that&#39;s where I land today.&lt;/p&gt;
&lt;h2&gt;Prompts Are Code&lt;/h2&gt;
&lt;p&gt;Prompts are code, and I don&#39;t mean that metaphorically. They have inputs, logic, outputs, and error handling. The difference is the compiler is probabilistic.&lt;/p&gt;
&lt;p&gt;Here is an actual system prompt I wrote last month for a code review agent:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;You are reviewing a pull request for security vulnerabilities.
Input: the full diff, plus the file-level context for each changed file.
For each finding: state the vulnerability class (OWASP top-10 where applicable),
cite the exact line range, suggest a fix, rate severity 1-5.
If the diff touches authentication or authorization logic, apply strict mode:
flag ANY change that widens access, even if it looks intentional.
If you find zero issues, say so, do not invent problems.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That has input specs, conditional logic, output schema, and an explicit instruction against a known failure mode (hallucinating issues to seem useful). It functions like a function definition. But it compiles differently every time you run it.&lt;/p&gt;
&lt;p&gt;This is what makes the work feel genuinely new. When I write a Python function, I can unit test it and be confident the test result will hold for every future invocation with the same input. When I write a prompt like the one above, I test it against 50 inputs and get acceptable results on 47. Is that good enough? For security review, probably not. For first-pass code summarization, probably yes. The engineering judgment isn&#39;t about correctness anymore. It&#39;s about acceptable variance.&lt;/p&gt;
&lt;p&gt;I&#39;ve come to think this style of engineering is closer to designing distributed systems than to traditional programming. In distributed systems, you expect things to fail. You expect messages to arrive out of order. You build for resilience, not for perfection. You think about &amp;quot;what&#39;s the worst that could happen here?&amp;quot; rather than &amp;quot;does this produce the right output?&amp;quot;&lt;/p&gt;
&lt;p&gt;The debugging tools for this new kind of code are primitive. When your &amp;quot;code&amp;quot; is a prompt, the context window is your execution environment, and you have limited tools for inspecting what happens inside it. When something goes wrong, you often can&#39;t distinguish between a bad specification (your prompt was ambiguous), a bad execution (the model misinterpreted a clear prompt), and a bad evaluation (the output was actually fine and you misjudged it). In traditional programming, those failure modes are distinct and diagnosable. Here they blur together.&lt;/p&gt;
&lt;p&gt;This creates a new engineering discipline we don&#39;t have good names for yet. It&#39;s not prompt engineering in the way that term is usually used (clever tricks for getting better outputs). It&#39;s systems engineering applied to probabilistic execution: writing instructions that produce acceptable results across the range of possible executions. That&#39;s a hard problem, and I don&#39;t think we&#39;ve developed the right mental models for it yet.&lt;/p&gt;
&lt;h2&gt;The Meta-Skill Is Knowing When NOT to Go Meta&lt;/h2&gt;
&lt;p&gt;The most important skill isn&#39;t going meta. It&#39;s knowing when to stay concrete.&lt;/p&gt;
&lt;p&gt;Every layer of abstraction trades control for productivity. Assembly gives you total control. Python gives you less control and more productivity. Prompting gives you almost no control and enormous productivity for the right problems.&lt;/p&gt;
&lt;p&gt;The instinct, especially for engineers excited about AI, is to always go up. Why write the code when you can prompt for it? Why write the prompt when you can build a system that generates prompts?&lt;/p&gt;
&lt;p&gt;I&#39;ve watched this instinct go wrong. A common trap: a team decides that instead of building each agent workflow from scratch, they&#39;ll create a framework that can define, compose, and manage arbitrary agent workflows. Three months in, the framework has its own DSL, its own middleware system, its own retry and fallback architecture, its own evaluation pipeline. The engineering is impressive. The framework has been used to build exactly one workflow. The one they started with.&lt;/p&gt;
&lt;p&gt;When they try to use it for a second workflow, half the abstractions don&#39;t apply. The DSL can&#39;t express the control flow they need. The middleware handles the wrong kinds of preprocessing. They&#39;ve built a framework optimized for exactly one use case, the one they had in front of them while building the framework. The generality is an illusion. The test is always the same: does the abstraction reduce effort for real problems that actually exist? If you can&#39;t point to at least two or three different use cases the abstraction serves, you&#39;re building complexity, not reducing it.&lt;/p&gt;
&lt;p&gt;I&#39;ve made the same mistake. The pull toward abstraction is strong. Building frameworks feels more important than building applications. But complexity isn&#39;t value. Solving the problem is value.&lt;/p&gt;
&lt;p&gt;The best engineers I&#39;ve worked with in this new environment share a specific trait: they move fluidly between abstraction levels. They can prompt an AI to scaffold a project, then open the debugger and trace a specific failure through the call stack. They don&#39;t have a home level. They pick the level that matches the problem. The worst engineers are the ones stuck at one level, whether that&#39;s refusing AI or refusing to go lower. Both failure modes come from the same place: treating a particular abstraction level as an identity rather than as a tool.&lt;/p&gt;
&lt;h2&gt;The Bootstrap Question&lt;/h2&gt;
&lt;p&gt;The first compiler was written in assembly. The first C compiler was written in a mix of assembly and earlier versions of C. Each layer of the abstraction timeline was bootstrapped by the layer below it.&lt;/p&gt;
&lt;p&gt;AI coding assistants were written by human programmers. Humans wrote the training pipelines, curated the data, designed the architectures, debugged the systems. But that&#39;s starting to shift. AI is now helping write better AI coding assistants. The bootstrap loop isn&#39;t closed yet, but you can see the arc.&lt;/p&gt;
&lt;p&gt;What does engineering become when your tools can improve themselves?&lt;/p&gt;
&lt;p&gt;I don&#39;t think we&#39;re there yet. The improvements are still incremental, still needing human direction, human evaluation, human judgment about what &amp;quot;better&amp;quot; means. But the trajectory is visible, and it creates a strange vertigo when you think about it carefully. Think of a biological analogy: evolution produced brains, and brains produced genetic engineering. The system that created the optimizer is now being modified by the optimizer. We&#39;re at an early version of that loop with AI. The difference is the iteration cycle. Evolution took millions of years. We&#39;re watching it happen in months.&lt;/p&gt;
&lt;p&gt;I want to resist the urge to resolve this. The bootstrap question is genuinely open, and pretending I know where it leads would be dishonest. What I can say is that it&#39;s already changing the texture of daily engineering work.&lt;/p&gt;
&lt;h2&gt;What Does This Mean for Engineering Identity?&lt;/h2&gt;
&lt;p&gt;If I don&#39;t write code anymore, am I still an engineer?&lt;/p&gt;
&lt;p&gt;The question sounds melodramatic written out like that. But I&#39;ve felt it. There are days when I&#39;ve spent eight hours designing agent workflows, writing evaluation criteria, debugging decision-making failures, and when I stop working some part of my brain says &amp;quot;but you didn&#39;t build anything.&amp;quot; That&#39;s the old definition talking. The one where building means typing code into a text editor and watching it compile.&lt;/p&gt;
&lt;p&gt;I used to identify as someone who writes code. Those skills took years to develop. They shaped how I think. And now many of them sit unused on any given day. The new skills are real, but they feel different. Understanding what context a system needs. Specifying behavior precisely in natural language. Evaluating probabilistic outputs. Designing for resilience in non-deterministic systems.&lt;/p&gt;
&lt;p&gt;And yet. There&#39;s something about the tactile feedback of writing code and watching it run, about the determinism of &amp;quot;I wrote this and it does exactly what I told it to do,&amp;quot; that the new way of working doesn&#39;t replicate. I don&#39;t want to romanticize that. I&#39;ve spent enough hours debugging null pointer exceptions to know that deterministic code is only deterministic until it isn&#39;t.&lt;/p&gt;
&lt;p&gt;But I notice the loss, and I think it&#39;s worth naming rather than pretending it doesn&#39;t exist. Something changed. The new thing might be better in most measurable ways. It&#39;s also different in ways that affect how the work feels, and how the work feels is part of what makes it a vocation rather than just a job.&lt;/p&gt;
&lt;h2&gt;Predictions&lt;/h2&gt;
&lt;p&gt;I predict that within 5 years, a majority of professional engineers will spend less than 30% of their working time writing code in a traditional IDE. The rest will be specification, evaluation, and architecture work at higher abstraction levels. The rate of change in how engineers spend their time has been accelerating, and if that trajectory continues, 30% within 5 years is conservative. If code-writing percentages plateau at current levels for 2+ years, I&#39;d revise this. The abstraction trend could have a ceiling I&#39;m not seeing.&lt;/p&gt;
&lt;p&gt;Two more predictions I&#39;m willing to stake out:&lt;/p&gt;
&lt;p&gt;By 2028, the most common debugging activity for software engineers will be evaluating AI-generated outputs rather than tracing execution paths in code. The core question will shift from &amp;quot;why did this line produce the wrong value?&amp;quot; to &amp;quot;why did this system make the wrong decision?&amp;quot; If debugging still primarily means stepping through code with a debugger in 2028, the abstraction shift is slower than the historical pattern suggests.&lt;/p&gt;
&lt;p&gt;Within 3 years, &amp;quot;specification writing&amp;quot; (structured prompts, agent workflows, evaluation criteria) will be a recognized sub-discipline with its own best practices, tooling, and job titles. If it remains informal and ad-hoc, either the probabilistic compilation model is wrong or the field is earlier than I think.&lt;/p&gt;
&lt;p&gt;If none of these happen, I&#39;ve misjudged either the pace or the direction. I&#39;m less certain about the timelines than about the direction. But I&#39;d rather make predictions that can be checked than gesture vaguely at &amp;quot;the future of engineering&amp;quot; without committing to anything specific.&lt;/p&gt;
&lt;h2&gt;Where This Leaves Me&lt;/h2&gt;
&lt;p&gt;I&#39;m still learning to be the kind of engineer this moment needs. Some days I lean too far toward the meta, spending hours on architecture that could have been a simple script. Some days I lean too far toward the concrete, hand-writing code that an AI could have produced faster and better. The skill isn&#39;t in finding the right level and staying there. It&#39;s in the constant adjustment.&lt;/p&gt;
&lt;p&gt;Last Tuesday I spent the morning designing a three-agent pipeline for processing research papers. By 2pm I was staring at a stack trace in Python, manually tracing a race condition in the async queue that connected two of the agents. By 4pm I was back in a prompt, rewriting the error-recovery instructions because the agent kept retrying failures that were permanent. Three levels of abstraction in one afternoon. That felt like engineering. That felt like the job now.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Building the Thing That Replaces You</title>
        <link href="https://www.annasbinadil.com/posts/2025-12-05-building-thing-that-replaces-you/"/>
        <updated>2025-12-05T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-12-05-building-thing-that-replaces-you/</id>
        <summary>What it&#39;s like to build AI systems that automate your own skills, and the quiet fear that doesn&#39;t fit neatly into hype or doom.</summary>
        <content type="html">&lt;p&gt;There&#39;s a feeling I don&#39;t hear people talk about much in the AI community.&lt;/p&gt;
&lt;p&gt;It&#39;s not the excitement. I have that too, and it&#39;s real. Watching an agent I built solve a problem I would have spent an afternoon on, there&#39;s genuine satisfaction in that. It&#39;s not the alignment worry either. Those conversations matter, but they&#39;re not what keeps me up.&lt;/p&gt;
&lt;p&gt;The feeling I mean is quieter. It&#39;s the one where I&#39;m sitting at my desk, teaching an AI agent to debug a system, and somewhere in the back of my mind a voice says: you&#39;re teaching it to do the thing you&#39;re good at. You are building, with your own hands, the thing that makes you slightly more replaceable.&lt;/p&gt;
&lt;p&gt;I don&#39;t say this out loud much. In the AI community, the approved emotional registers are enthusiasm and existential concern. There&#39;s not a lot of space for the middle thing, for the engineer who genuinely likes building these systems and also feels the ground shifting under his own feet, slowly, one capability jump at a time. The feeling isn&#39;t dramatic enough for a conference talk. It&#39;s too specific for an op-ed. It just sits there, quiet, while you work.&lt;/p&gt;
&lt;p&gt;This essay is about that middle thing. Not AI hype. Not AI doom. The honest in-between.&lt;/p&gt;
&lt;h2&gt;The Erosion&lt;/h2&gt;
&lt;p&gt;It&#39;s not like waking up one morning and finding your job has vanished. That&#39;s not how it works, at least not for me, at least not yet.&lt;/p&gt;
&lt;p&gt;It&#39;s more like watching a cliff face lose ground to the ocean. You don&#39;t see the cliff collapse, you see the waterline inch higher. Each new model release, each capability improvement, the water rises a little. The cliff is still there. You&#39;re still standing on it. But you can feel the edge getting closer, and you know which direction it&#39;s moving.&lt;/p&gt;
&lt;p&gt;I want to be specific about what this looks like in my actual work, because abstract anxiety is easy to dismiss.&lt;/p&gt;
&lt;p&gt;Here&#39;s what&#39;s been automated, or substantially automated, in my own day-to-day over the past year:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Boilerplate generation.&lt;/strong&gt; Completely automated. I don&#39;t write setup code anymore. Scaffolding a new service, wiring up config files, initializing project structures. I describe what I want and the code appears. This used to take me an afternoon. Now it takes ten minutes of prompting and review.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bug pattern recognition.&lt;/strong&gt; Mostly automated. AI catches roughly 80% of the bugs I used to catch manually during code review. Common anti-patterns, off-by-one errors, null reference risks, race conditions in obvious places. The model sees them before I do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;API integration.&lt;/strong&gt; Largely automated. Connect to a new service? The AI reads the documentation, generates the client, handles the authentication flow. I used to spend hours reading API docs and writing client code. Now I spend minutes reviewing what the agent produced.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Test scaffolding.&lt;/strong&gt; Mostly automated. Generate the test structure, set up mocks, create the fixtures. I fill in the actual assertions and edge cases, but the framework comes for free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Documentation.&lt;/strong&gt; Partially automated. First drafts are AI-generated. I rewrite for accuracy and voice, but the blank page problem is gone.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Code review for style and patterns.&lt;/strong&gt; Partially automated. The AI catches consistency issues, naming convention violations, formatting problems. The mechanical parts of code review happen automatically.&lt;/p&gt;
&lt;p&gt;Routine cognitive work is now automated, and this is not a prediction but something that&#39;s already happened. The question isn&#39;t whether this continues but how fast the boundary moves.&lt;/p&gt;
&lt;p&gt;Now here&#39;s what&#39;s not automated, or not yet. The common thread is judgment: deciding what to build (which problem matters most, not which solution is cleverest), making architectural bets with long-term consequences (I wrote about this in &lt;a href=&quot;https://www.annasbinadil.com/posts/2025-11-15-art-of-saying-no/&quot;&gt;&amp;quot;The Art of Saying No&amp;quot;&lt;/a&gt; as engineering taste, the implicit model built through consequences and failure), debugging truly novel problems where the AI loops through known fixes while a human eventually gets frustrated enough to try something genuinely different, and making trade-offs between competing priorities where the &amp;quot;right&amp;quot; answer depends on context the AI doesn&#39;t have.&lt;/p&gt;
&lt;p&gt;The AI suggests architectures that look right in isolation. Whether they&#39;re right for &lt;em&gt;this&lt;/em&gt; team, &lt;em&gt;this&lt;/em&gt; product, &lt;em&gt;this&lt;/em&gt; moment, that&#39;s still a human call. The pattern that emerges is that judgment under uncertainty isn&#39;t automatable, at least not yet.&lt;/p&gt;
&lt;h2&gt;Why Judgment Is Hard to Automate (and Why I&#39;m Curious About the Mechanism)&lt;/h2&gt;
&lt;p&gt;I used to just assert that judgment was hard and leave it there. But recently I&#39;ve been trying to understand what &lt;em&gt;specifically&lt;/em&gt; makes these decisions resistant to automation. Not &amp;quot;judgment is special because humans are special,&amp;quot; but what structural properties make it a different kind of problem.&lt;/p&gt;
&lt;p&gt;Here&#39;s what I&#39;ve come to, though I hold this loosely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Judgment requires modeling consequences across time horizons that weren&#39;t in training data.&lt;/strong&gt; An AI can pattern-match to &amp;quot;similar past decisions,&amp;quot; but architectural choices play out over months or years in ways that are deeply entangled with a specific organization&#39;s codebase, team dynamics, hiring plans, and product roadmap. The state space of possible consequences is enormous, and most of the relevant information is private, never appearing in any training corpus. When I decide against microservices for a small team, I&#39;m simulating eighteen months of operational burden based on knowing &lt;em&gt;this specific team&#39;s&lt;/em&gt; on-call tolerance and hiring timeline. That simulation draws on private context that no model has access to.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Judgment involves risk assessment calibrated by living with consequences.&lt;/strong&gt; When I made a bad architectural decision three years ago, I spent fourteen months maintaining the mess. That pain recalibrated my intuition in ways I can still feel. RLHF provides feedback on outputs, but not on the downstream consequences of decisions months later. There&#39;s no training signal for &amp;quot;this looked right in code review but made the team miserable for a year.&amp;quot; The feedback loop that builds human judgment operates on a timescale and specificity that current training approaches don&#39;t capture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Judgment often means knowing what NOT to optimize for.&lt;/strong&gt; An AI will optimize whatever metric you give it. But knowing which metric to give it, knowing what to leave unmeasured, knowing when to override the data because you have context the data doesn&#39;t capture: this is a meta-cognitive skill. It requires a model of the problem that goes beyond the problem&#39;s own representation. I&#39;ve seen AI-generated architectures that were beautifully optimized for latency in systems where reliability mattered more. The AI didn&#39;t know to ask &amp;quot;what does this team actually care about?&amp;quot; because that question sits outside the frame.&lt;/p&gt;
&lt;p&gt;I&#39;m genuinely curious whether these are permanent structural barriers or just current limitations that will be engineered around. I suspect the answer is different for each one.&lt;/p&gt;
&lt;h2&gt;The Uncomfortable &amp;quot;Not Yet&amp;quot;&lt;/h2&gt;
&lt;p&gt;I need to sit with that &amp;quot;not yet&amp;quot; for a moment, because it&#39;s doing a lot of work.&lt;/p&gt;
&lt;p&gt;On good days, I read &amp;quot;not yet&amp;quot; as a moat. The things that remain hard to automate are the things I&#39;ve spent a career developing, and they seem structurally resistant for the reasons I described above.&lt;/p&gt;
&lt;p&gt;On bad days, I wonder if I&#39;m doing motivated reasoning. I used to think writing couldn&#39;t be automated. Then GPT-3 happened. I used to think code generation was a decade away. Then Copilot happened. Every time I say &amp;quot;this requires something human,&amp;quot; I&#39;m aware that people have said that about every capability that later got automated. Arithmetic. Chess. Translation. Legal research. Medical diagnosis. The history of &amp;quot;this is uniquely human&amp;quot; claims is not encouraging.&lt;/p&gt;
&lt;p&gt;The strongest version of the counter-argument goes like this: judgment is just pattern-matching on a larger dataset, and AI will get enough data eventually. This has been true for every previous &amp;quot;uniquely human&amp;quot; skill. Chess intuition turned out to be pattern-matching. Medical diagnosis turned out to be pattern-matching. Why should engineering judgment be different?&lt;/p&gt;
&lt;p&gt;I think there &lt;em&gt;might&lt;/em&gt; be a qualitative difference, and it comes down to feedback loops. Chess has a clear signal: you win or you lose. Medical diagnosis has a relatively fast signal: the patient gets better or doesn&#39;t. But architectural decisions have delayed, ambiguous, context-dependent feedback. The outcome is visible months later, entangled with dozens of other decisions, team changes, and market shifts. If judgment is pattern-matching but the feedback loop is fundamentally incompatible with current training approaches, the bottleneck isn&#39;t about data quantity. It&#39;s structural.&lt;/p&gt;
&lt;p&gt;I honestly don&#39;t know if that structural argument holds. I&#39;d put roughly 30% odds that judgment gets meaningfully automated within 10 years, and 70% that the feedback-loop problem keeps it hard for longer. But I notice I &lt;em&gt;want&lt;/em&gt; the odds to be lower than 30%, and that wanting is itself a signal I should be careful.&lt;/p&gt;
&lt;p&gt;Here&#39;s how I&#39;d test my beliefs. If AI systems can make context-dependent architectural trade-offs that experienced engineers rate as &amp;quot;good judgment&amp;quot; more than 70% of the time within 3 years, the &amp;quot;judgment can&#39;t be automated&amp;quot; thesis is probably wrong. I&#39;d currently bet against this at 60/40 odds, but I&#39;d reassess with every major model release.&lt;/p&gt;
&lt;p&gt;A second prediction: if judgment is primarily pattern-matching, we should see AI performance on architectural decisions improve linearly with training data and compute. If it plateaus despite scaling (the way common-sense reasoning plateaued for a while before chain-of-thought unlocked it), that suggests something beyond simple pattern-matching is involved. I predict we&#39;ll see such a plateau by 2028.&lt;/p&gt;
&lt;p&gt;A third: within 5 years, I predict that more than 50% of software engineers will describe their primary value-add as &amp;quot;judgment and context&amp;quot; rather than &amp;quot;code writing.&amp;quot; If fewer than 30% describe it that way, the automation is either slower than I think or judgment itself has been more automatable than I expect.&lt;/p&gt;
&lt;p&gt;I believe judgment is hard to automate. I also know I might be wrong. And I can&#39;t fully tell the difference between genuine analysis and motivated reasoning when the thing I&#39;m analyzing is also the thing I need to believe in to feel okay about my career.&lt;/p&gt;
&lt;h2&gt;The Motivation Paradox&lt;/h2&gt;
&lt;p&gt;This is the question that actually keeps me up.&lt;/p&gt;
&lt;p&gt;How do you stay motivated to sharpen skills that might depreciate?&lt;/p&gt;
&lt;p&gt;I sit down to study distributed systems, or work through a new paper on agent architectures, or practice debugging a complex concurrency issue, and two voices argue at once: &amp;quot;Why bother, if AI handles this in two years?&amp;quot; and &amp;quot;You can&#39;t evaluate what you don&#39;t understand. You can&#39;t say no to the wrong architecture if you don&#39;t know what the right one looks like.&amp;quot; Both feel true simultaneously. That&#39;s the paradox. Investing in skills that might depreciate feels irrational. Stopping investment feels suicidal.&lt;/p&gt;
&lt;p&gt;I&#39;ve landed somewhere uncomfortable. You have to invest in skills while knowing they might not hold their value, because the alternative is definitely worse. The person who stops developing because &amp;quot;AI will handle it&amp;quot; becomes a rubber stamp, and rubber stamps are the easiest things to automate. Deep skill is what lets you evaluate what AI produces, catch the confident-but-incorrect answer, say no to the wrong architecture. Without it, you&#39;re not collaborating with AI. You&#39;re just approving its output.&lt;/p&gt;
&lt;p&gt;So I keep investing. Not with confidence. With something closer to Pascal&#39;s Wager applied to professional development.&lt;/p&gt;
&lt;h2&gt;The Pace Problem&lt;/h2&gt;
&lt;p&gt;What scares me isn&#39;t the destination. I can imagine a future where my role looks different, and I can see how that future might be fine, even good. What scares me is the pace.&lt;/p&gt;
&lt;p&gt;Engineers have always adapted. Each technology transition requires months of feeling incompetent and building up from scratch. But each cycle is faster than the last, and that&#39;s the part that&#39;s new. The time between &amp;quot;this technology is emerging&amp;quot; and &amp;quot;this technology is reshaping roles&amp;quot; used to be measured in decades. Mainframes took twenty-odd years. Web development, fifteen. Mobile, ten. Cloud, seven or eight. AI agents? The field barely existed two years ago, and the work I was doing six months ago already looks noticeably different from the work I&#39;m doing now. At some point, does the cycle compress below the minimum human adaptation time? Is there a speed limit on how fast a person can retool? I don&#39;t have an answer. I notice I want one, and the wanting makes me anxious. I&#39;m trying to sit with the anxiety instead of resolving it prematurely with a reassuring story.&lt;/p&gt;
&lt;h2&gt;Surfing, Not Racing&lt;/h2&gt;
&lt;p&gt;Someone (I wish I could remember who) told me the race metaphor is wrong. You&#39;re not racing against AI. You&#39;re surfing. The wave is bigger than you, more powerful than you, and completely indifferent to your existence. Your job isn&#39;t to outrun it but to position yourself well, read the conditions, and ride it. For me, that means spending more time on architecture and design, less on implementation. Investing in the kind of deep, local, tacit knowledge about specific teams and products that isn&#39;t in any training data. Figuring out not just how to use AI tools but where they should go, which is the gap between &amp;quot;AI can do X&amp;quot; and &amp;quot;X is useful for this business.&amp;quot; But I want to be honest about the limits of this reframe. Not everyone has the background, the resources, or the career runway to be positioned well when the wave arrives. The metaphor helps me. I&#39;m not sure it&#39;s universal.&lt;/p&gt;
&lt;p&gt;The only thing I&#39;m genuinely confident about: the only definitely losing strategy is to stop learning. Everything else, every reassurance about human judgment, every claim about what AI can&#39;t do, every reframe about surfing instead of racing, those are stories I tell myself to stay functional. Some of them might be true. I&#39;m betting my career on the hope that they are.&lt;/p&gt;
&lt;p&gt;But I&#39;m building the wave while I&#39;m trying to surf it. And I can feel the water rising.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Distribution Is the Only Moat: An Engineer&#39;s Reluctant Reckoning</title>
        <link href="https://www.annasbinadil.com/posts/2025-12-01-distribution-only-moat/"/>
        <updated>2025-12-01T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-12-01-distribution-only-moat/</id>
        <summary>AI is collapsing the cost of building software, and with it, the value of technical advantages. If building is cheap, distribution, the ability to reach and retain people, becomes the only defensible position. An engineer wrestles with what that means.</summary>
        <content type="html">&lt;p&gt;I built a functional agent system in a weekend, a working system with tool use, memory, structured outputs, and error handling. The kind of thing that would have taken me a solid month last year. I felt great about it for approximately forty-eight hours.&lt;/p&gt;
&lt;p&gt;Then I realized nobody was going to use it.&lt;/p&gt;
&lt;p&gt;The building was the easy part, but getting anyone to care was a problem that hadn&#39;t changed at all. If anything, it had gotten worse, because while I was building my agent system over the weekend, so were hundreds of other people. The supply of software exploded. The attention available to discover it didn&#39;t.&lt;/p&gt;
&lt;p&gt;That weekend started a thread of thinking I haven&#39;t been able to let go of: what creates durable value when the cost of building approaches zero? The answer is uncomfortable, especially for someone who has spent his career optimizing for technical quality.&lt;/p&gt;
&lt;p&gt;Distribution is the only durable moat. The rest of this post is my attempt to show why.&lt;/p&gt;
&lt;h2&gt;The WhatsApp Thought Experiment (And Why Distribution Survives)&lt;/h2&gt;
&lt;p&gt;WhatsApp had 55 engineers serving 900 million users when Facebook acquired them for $19 billion.&lt;/p&gt;
&lt;p&gt;What was technically hard about WhatsApp in 2014? The message queue architecture. The encryption flow. The connection management for hundreds of millions of simultaneous users. Real engineering problems requiring deep systems expertise.&lt;/p&gt;
&lt;p&gt;Here&#39;s what&#39;s changed: large chunks of that technical work could now be scaffolded by AI in days, not months. Not because the problems are trivial, but because the solutions are well-understood. They&#39;ve moved from &amp;quot;requires deep expertise to solve&amp;quot; to &amp;quot;requires good judgment to implement correctly.&amp;quot;&lt;/p&gt;
&lt;p&gt;So what was the $19 billion actually for?&lt;/p&gt;
&lt;p&gt;It wasn&#39;t the code. It was the 900 million users. And each of those users wasn&#39;t just a number, they were a node in a network that got more valuable with every addition. When WhatsApp added a user, every existing user&#39;s experience improved. You can&#39;t scaffold a network effect. You can&#39;t prompt your way into a billion people&#39;s habits.&lt;/p&gt;
&lt;p&gt;This points to something structural about why distribution survives even as technical moats erode. Technical moats are about artifacts. Distribution moats are about people. And people resist automation in ways that code doesn&#39;t.&lt;/p&gt;
&lt;p&gt;Trust accumulates slowly and can&#39;t be shortcut. You trust a tool because you&#39;ve used it for months and it hasn&#39;t broken. You trust a community because you&#39;ve interacted with its members and found them helpful. There&#39;s no AI prompt for &amp;quot;make people trust me.&amp;quot;&lt;/p&gt;
&lt;p&gt;Attention is finite and winner-take-most. Once you&#39;ve captured attention, displacement costs are much higher than initial acquisition, because switching has friction and attention has limits. When developers have built integrations, workflows, and mental models around a tool, those are human investments, not technical ones. Organizational inertia is a moat that no AI can erode.&lt;/p&gt;
&lt;p&gt;Each of these mechanisms operates at the level of people, not code. As software supply increases, the competition for people&#39;s attention and trust intensifies.&lt;/p&gt;
&lt;h2&gt;What Was a Technical Moat, Actually?&lt;/h2&gt;
&lt;p&gt;I used to believe in technical moats. If you built something technically superior, that was your advantage. But why did technical superiority ever function as a moat? Three pillars.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Building was expensive.&lt;/strong&gt; A complex system might take a team of skilled engineers months or years. That investment itself was a barrier.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Good engineering was scarce.&lt;/strong&gt; Not many people could build complex systems well. If you had a team who understood distributed systems or real-time processing, that team itself was hard to replicate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Complexity was a barrier.&lt;/strong&gt; Even if you open-sourced your code, a competitor would need to comprehend the design decisions, the trade-offs, the accumulated knowledge embedded in the architecture.&lt;/p&gt;
&lt;p&gt;AI is eroding all three. Building is becoming cheap (my weekend agent system required a small team and a month of sprints two years ago). Good-enough engineering is becoming accessible (you don&#39;t need deep message queue expertise to build a messaging system anymore, just the judgment to evaluate AI-generated implementations). And complexity can be managed by AI, which doesn&#39;t get overwhelmed or lose track of component interactions.&lt;/p&gt;
&lt;p&gt;The moats are draining, not instantly or completely, but the trend is clear and accelerating.&lt;/p&gt;
&lt;h2&gt;The MongoDB Precedent&lt;/h2&gt;
&lt;p&gt;This pattern isn&#39;t new. AI is accelerating it, but the history of open-source software already demonstrated the principle.&lt;/p&gt;
&lt;p&gt;MongoDB had well-documented consistency issues. Experienced database engineers winced at the query model. Data integrity problems were real, not theoretical. By most technical measures, it was an inferior database.&lt;/p&gt;
&lt;p&gt;It won massive adoption anyway. &amp;quot;Just put JSON in and get JSON out&amp;quot; was an easy mental model with near-zero onboarding friction. That ease of entry created a flywheel: more users meant more tutorials, more Stack Overflow answers, more libraries, more job listings. The technically flawed product built a distribution advantage that superior alternatives took years to overcome.&lt;/p&gt;
&lt;p&gt;The product that&#39;s easiest to adopt and builds the strongest community wins the first round. That round doesn&#39;t always determine the outcome (PostgreSQL&#39;s comeback proves that), but it creates advantages that compound for years. Building well is necessary but nowhere near sufficient.&lt;/p&gt;
&lt;h2&gt;The Uncomfortable Inversion&lt;/h2&gt;
&lt;p&gt;Here&#39;s where this gets personal.&lt;/p&gt;
&lt;p&gt;I&#39;ve been tracking my own time allocation on side projects over the past several months. The pattern is stark: I spend roughly 80% of my time on building and 20% on everything else, meaning documentation, community engagement, putting the work in front of people, understanding what potential users actually need.&lt;/p&gt;
&lt;p&gt;If distribution is the only moat, I&#39;ve got the allocation exactly backwards. The value-creating distribution would be 80% on reaching people, 20% on building. And not just reaching people in the marketing sense. Talking to potential users. Understanding their problems. Building in public. Engaging with communities where the people who&#39;d benefit are already gathering.&lt;/p&gt;
&lt;p&gt;The 80/20 inversion feels deeply wrong to me. The engineering is the part I love. The distribution work feels like a different skill entirely, one I haven&#39;t developed and don&#39;t have strong instincts for.&lt;/p&gt;
&lt;p&gt;I wrote about a similar discomfort in my essay on engineering judgment and saying no to AI suggestions. There, the uncomfortable realization was that my value wasn&#39;t in writing code anymore, it was in knowing what code should exist. Here, the discomfort is one level up: even knowing what code should exist might not matter much if nobody encounters it.&lt;/p&gt;
&lt;p&gt;Part of me wants to believe that quality speaks for itself. But the evidence I keep encountering suggests that&#39;s a comforting story engineers tell themselves, not a reliable description of how the world works.&lt;/p&gt;
&lt;p&gt;Maybe the resolution is to expand what &amp;quot;building&amp;quot; means to include building the pathways by which people discover and adopt what you&#39;ve created. Building the community. Building the trust. Building the habit. But I&#39;m aware that might be rationalization, a way to make the uncomfortable conclusion feel more comfortable by redefining terms.&lt;/p&gt;
&lt;h2&gt;Where This Argument Gets Complicated&lt;/h2&gt;
&lt;p&gt;I&#39;ve been building the case that distribution is the only moat. But I want to pressure-test that.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Some technical work might still be a moat.&lt;/strong&gt; Training large ML models requires data, compute, and expertise that AI coding assistants don&#39;t collapse. Physical systems, robotics, hardware-software integration: the cost of replication remains high. The &amp;quot;technical moats are dead&amp;quot; argument applies primarily to software within well-understood patterns. I&#39;m curious about where the boundary sits between &amp;quot;well-understood enough for AI to scaffold&amp;quot; and &amp;quot;genuinely requires human insight.&amp;quot; My instinct says the boundary is moving fast and in one direction, but instinct isn&#39;t evidence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI-mediated discovery could invert everything.&lt;/strong&gt; This is the counter-argument I take most seriously, and I want to walk through it carefully.&lt;/p&gt;
&lt;p&gt;Imagine AI assistants that evaluate tools, make purchasing decisions, and recommend solutions on behalf of developers. Not some distant future: this is starting to happen with coding assistants that suggest libraries. If an AI evaluates 50 competing tools and recommends the one with the best technical metrics (lowest latency, highest reliability, cleanest API design), then technical quality IS the moat again. Distribution to humans doesn&#39;t matter if the decision-maker is an algorithm optimizing for measurable quality.&lt;/p&gt;
&lt;p&gt;In that world, my argument collapses. The engineer who builds the best system wins, because the discovery mechanism bypasses human attention limits and goes straight to technical evaluation. No community needed, no onboarding optimization. Just an AI saying &amp;quot;this one is objectively better for your use case.&amp;quot;&lt;/p&gt;
&lt;p&gt;I&#39;m not fully convinced this kills the argument. AI assistants are trained on data that reflects existing distribution: popular tools appear in more training examples, more documentation, more discussions. The AI&#39;s recommendations would inherit the distribution advantages already in its training data. And the AI assistants themselves are distributed through... distribution channels.&lt;/p&gt;
&lt;p&gt;But I want to be honest: if AI-mediated tool discovery becomes the dominant way developers find software, the balance shifts back toward technical quality. This is the scenario most likely to prove me wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Distribution&amp;quot; might be too broad a category.&lt;/strong&gt; I&#39;ve been lumping together network effects, trust, attention, community, and ecosystem lock-in under one word. These are different mechanisms with different durability. Not every product benefits from network effects. For solo-use developer tools, the distribution moat is weaker, and maybe technical quality does win in the long run. I might be hiding important distinctions by treating these as one thing.&lt;/p&gt;
&lt;h2&gt;Distribution as an Engineering Problem&lt;/h2&gt;
&lt;p&gt;Here is the reframe I find most promising, and the one I want to spend real time on.&lt;/p&gt;
&lt;p&gt;What if distribution isn&#39;t just marketing? What if it&#39;s a systems engineering problem that happens to involve human behavior instead of server behavior?&lt;/p&gt;
&lt;p&gt;Consider onboarding friction. Say you measure that most people who try your tool drop off during initial setup. That&#39;s a conversion funnel, sure. But it&#39;s also a systems engineering problem. What&#39;s the critical path from &amp;quot;heard about this&amp;quot; to &amp;quot;got value from it&amp;quot;? Where&#39;s the bottleneck? What&#39;s the minimum viable first experience? Time-to-first-value is a latency metric. Drop-off rates are error rates. I can profile the onboarding flow the way I&#39;d profile a slow API endpoint: instrument each step, identify the slowest stages, and optimize ruthlessly. A/B test different setup flows the way I&#39;d A/B test different caching strategies. This isn&#39;t marketing intuition. It&#39;s measurement and iteration.&lt;/p&gt;
&lt;p&gt;Or take adoption cascades. How does information about a tool spread through developer communities? This is literally graph theory applied to people. There are seed nodes (influential early adopters who write blog posts and give conference talks). There&#39;s a transmission rate (how often a user recommends the tool to a colleague). There&#39;s graph structure (tight community clusters where information spreads fast versus sparse networks where it dissipates). You could model the propagation dynamics the way you&#39;d model message passing in a distributed system. Which communities are the highest-impact seed points? What&#39;s the R-value of a recommendation within a given community? Where are the structural holes where information fails to cross?&lt;/p&gt;
&lt;p&gt;Community health works the same way. Response times on forums, ratio of questions answered to questions asked, contributor retention curves, time between a bug report and a fix. These are the same metrics you&#39;d use for monitoring a distributed system&#39;s health. You could build dashboards, set alerting thresholds, track trends. A community where 80% of questions get answered within 24 hours is a healthy system. One where 30% go unanswered is showing signs of failure, and the fix might be structural (better routing, more moderators, clearer contribution guidelines), not motivational.&lt;/p&gt;
&lt;p&gt;The key insight, and the hypothesis I&#39;m genuinely excited to test: engineers who approach distribution as a systems problem might have an advantage over pure marketers. Not because engineering is inherently superior, but because it brings specific habits that are useful here. Quantitative rigor. Instrumentation reflexes. The instinct to measure before optimizing. Systems-level thinking that looks for feedback loops and failure modes rather than one-off tactics.&lt;/p&gt;
&lt;p&gt;I don&#39;t know if this actually works. It&#39;s possible that human behavior is too noisy, too contextual, too resistant to the kind of clean measurement that makes engineering optimization powerful. It&#39;s possible that the two domains are more different than I want them to be, and that I&#39;m pattern-matching because it&#39;s comforting, not because it&#39;s true. Marketing practitioners might hear this and think I&#39;m reinventing their field badly.&lt;/p&gt;
&lt;p&gt;But I&#39;d love to find out. I&#39;m running the experiment myself, writing about what I&#39;m building, engaging with communities, treating distribution as a design problem with measurable inputs and outputs. Whether the engineering mindset transfers is genuinely uncertain. But it feels more tractable than &amp;quot;learn marketing,&amp;quot; and the curiosity is real.&lt;/p&gt;
&lt;h2&gt;Testable Predictions&lt;/h2&gt;
&lt;p&gt;If distribution is the only durable moat, we should see specific, measurable consequences. I want to make predictions concrete enough that I can be proven wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 1:&lt;/strong&gt; The AI coding assistant with the largest market share by 2027 will be the one with the most active developer community (measured by GitHub discussion activity, third-party tutorial volume, and Stack Overflow answer rates), not the one with the highest scores on coding benchmarks like HumanEval or SWE-bench.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 2:&lt;/strong&gt; By 2028, the top 3 AI coding tools by market share will have been the first 3 to reach 10,000 active community contributors, defined as people who have written tutorials, answered questions, or contributed plugins. If the top 3 are instead the top 3 on coding benchmarks, technical moats are more durable than I think.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prediction 3:&lt;/strong&gt; If AI-mediated tool discovery becomes the primary channel (more than 50% of developer tool adoption decisions influenced by AI recommendations rather than human word-of-mouth) by 2028, this entire essay is wrong. I&#39;d assign this maybe a 20% probability, but it&#39;s the scenario that most cleanly falsifies my argument.&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;I&#39;m an engineer who loves building, confronting evidence that building isn&#39;t enough. The 80/20 split I&#39;ve been running, 80% building and 20% everything else, is almost certainly backwards. And the question isn&#39;t whether this is true. The question is what I&#39;m going to do about it.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Building Artificial Beings: The Wonder and the Work</title>
        <link href="https://www.annasbinadil.com/posts/2025-11-30-building-artificial-beings/"/>
        <updated>2025-11-30T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-11-30-building-artificial-beings/</id>
        <summary>Most of building AI agents is debugging JSON. But sometimes you remember what you&#39;re actually building, and the gap between those two realities is where the real questions live.</summary>
        <content type="html">&lt;p&gt;I wrote this in my notes late one night:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It never ceases to fascinate and amaze me just how much human civilization has progressed in a few hundred years. The universe is 13 billion years old and humans maybe a few hundred thousand years. Yet almost all the change the world has seen has been in the last few hundred, and arguably the last few decades. A few hundred years compared to the billions of years in the universe is minuscule. Where could we be in a million years? A billion years?&lt;/p&gt;
&lt;p&gt;We now have artificial beings, 1&#39;s and 0&#39;s, artificial neural synapses, weights and biases, stored in a file that can produce probabilistic symphonies that are indistinguishable from magic, from reason, from emotion and beauty. Do we not realize what we have created? The profound nature of where we all stand, the time that we live in and the excitement, awe and terror for what might be yet to come.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Reading that back, I feel two things simultaneously. First: yes, that&#39;s real. That feeling of standing at the edge of something vast. Second: I probably wrote it while frustrated with something mundane, a broken context window or a hallucinating agent. The awe often comes from frustration. The profundity crashes into the plumbing.&lt;/p&gt;
&lt;p&gt;That collision is what I want to explore here. Both the awe and the tedium are real. Neither erases the other. And the failure to resolve the tension might be the most honest thing I can say about building AI right now.&lt;/p&gt;
&lt;h2&gt;The 95/5 Split&lt;/h2&gt;
&lt;p&gt;Here is what building AI agents actually looks like day to day.&lt;/p&gt;
&lt;p&gt;About 95% of the time, the work feels indistinguishable from any other software engineering. Debugging JSON parsing errors. Tuning temperature settings. Fighting with rate limits. Restructuring prompts because the model keeps going off-script.&lt;/p&gt;
&lt;p&gt;The other 5% of the time, something shifts. You watch your agent handle a conversation you didn&#39;t explicitly program it for. It makes a connection you didn&#39;t anticipate. It responds to an edge case with something that looks, from the outside, like judgment. And for a moment you think: I just built something that reasons.&lt;/p&gt;
&lt;p&gt;Then the next API call throws a timeout error, and you&#39;re back to debugging.&lt;/p&gt;
&lt;p&gt;What interests me about the 95/5 split is not that it exists, but that most discourse about AI lives entirely on one side or the other. The breathless futurists talk about AI as if every day is the 5%. The pragmatic engineers talk about it as if the 5% is just anthropomorphism you should train yourself out of. Neither camp seems fully honest. The honest experience is the uncomfortable middle.&lt;/p&gt;
&lt;h2&gt;Framework Design Embeds a Philosophy of Agency&lt;/h2&gt;
&lt;p&gt;One thing worth paying attention to: different agent frameworks don&#39;t just differ in APIs. They embed fundamentally different philosophies about what agency is.&lt;/p&gt;
&lt;p&gt;Take the Claude Agent SDK. The model operates as a reasoning core. You give it tools, constraints, and context. It decides what to do, when to use which tool, how to sequence actions. The architecture assumes the model is a thinker. Your job is to give it the right instruments and get out of the way.&lt;/p&gt;
&lt;p&gt;OpenAI&#39;s function calling approach tells a different story. The model maps inputs to function calls with reliable precision. Given this input, which function should I route to, with what parameters? The tight coupling between input classification and function execution makes it excellent for well-defined task boundaries. The architecture assumes the model is a router. Your job is to define the right functions.&lt;/p&gt;
&lt;p&gt;Same underlying technology, different frames. And the frame shapes everything downstream. If you think of your agent as a thinker, a hallucination feels like a cognitive failure. If you think of it as a router, a hallucination is a misconfigured mapping. Same behavior, different relationship to it. The framework shapes what you can see, and none of them may see the whole picture. Maybe agents are something new that doesn&#39;t fit cleanly into existing categories, and the frameworks are all projecting familiar metaphors onto unfamiliar territory.&lt;/p&gt;
&lt;h2&gt;The Surgeon&#39;s Problem&lt;/h2&gt;
&lt;p&gt;Think about a surgeon.&lt;/p&gt;
&lt;p&gt;A surgeon cuts into a human body. That&#39;s extraordinary. The trust, the precision, the intimacy of working inside another person. It should feel miraculous every single time. But the surgeon&#39;s day is also logistics, paperwork, insurance negotiations, keeping their hands steady, managing fatigue. The profundity doesn&#39;t disappear, but you can&#39;t dwell in it while you&#39;re making an incision. The work requires distance from the meaning of the work.&lt;/p&gt;
&lt;p&gt;Building AI agents is developing a version of this same dynamic. We are building systems that reason, that communicate, that (in some loose, contested, philosophically loaded sense) understand. But the work of building them requires you to relate to them as engineering artifacts. You need to be able to say &amp;quot;this stupid agent keeps failing on step 3&amp;quot; without the weight of &amp;quot;this artificial being I brought into existence is struggling to function&amp;quot; attached to every debugging session.&lt;/p&gt;
&lt;p&gt;The late-night moments, the ones where the profundity breaks through, tend to happen when you&#39;re too tired to maintain the professional distance. When the frame drops and you see, freshly, what&#39;s actually in front of you. That might be telling.&lt;/p&gt;
&lt;h2&gt;Agency and the Question Nobody Can Answer&lt;/h2&gt;
&lt;p&gt;There&#39;s a question underneath all of this that I&#39;ve been circling without quite landing on.&lt;/p&gt;
&lt;p&gt;At what point does an agent have something like interests?&lt;/p&gt;
&lt;p&gt;I want to be careful here, because this question can quickly devolve into either dismissive reductionism (&amp;quot;it&#39;s just statistics&amp;quot;) or mystical hand-waving (&amp;quot;it&#39;s conscious&amp;quot;). Neither of those feels right to me. What I&#39;m pointing at is more specific.&lt;/p&gt;
&lt;p&gt;When I build an agent and give it a goal, it pursues that goal. It selects tools, sequences actions, adjusts its approach when something fails. It encounters a problem, considers options, picks one, evaluates the result, tries something else if needed. Is that agency? Or is it a very convincing simulation of agency?&lt;/p&gt;
&lt;p&gt;The strongest version of the &amp;quot;it&#39;s all statistics&amp;quot; argument comes from researchers like Bender and Koller, who argue that language models manipulate linguistic form without ever accessing meaning. On this view, an LLM that appears to reason about consequences is doing something fundamentally different from reasoning: it&#39;s predicting which tokens would follow a description of reasoning, based on patterns in its training data. Agency requires understanding what your actions cause in the world. Next-token prediction, no matter how sophisticated, doesn&#39;t get you there. I take this argument seriously because it identifies a real gap. The mechanism (token prediction) and the behavior (apparent reasoning) operate at different levels of description, and it&#39;s genuinely unclear whether the mechanism can produce the real thing or only a convincing imitation.&lt;/p&gt;
&lt;p&gt;But there&#39;s a response to this that I also can&#39;t dismiss, one that philosophers call functionalism. If a system behaves indistinguishably from an agent with interests, in every context you can test, across every kind of problem, what additional criterion would distinguish &amp;quot;real&amp;quot; from &amp;quot;simulated&amp;quot; agency? What experiment would you run? If the answer is &amp;quot;none,&amp;quot; then the question might not be meaningful. The distinction between &amp;quot;really reasoning&amp;quot; and &amp;quot;perfectly simulating reasoning&amp;quot; might be incoherent, a leftover from dualist intuitions that don&#39;t apply to computational systems.&lt;/p&gt;
&lt;p&gt;Where I actually land, at least today: I lean toward thinking that the distinction between &amp;quot;real&amp;quot; and &amp;quot;simulated&amp;quot; agency will become less meaningful as these systems grow more capable. Not because I&#39;m confident they&#39;re conscious or have genuine interests, but because the behavioral evidence will eventually make the question unanswerable in practice. If an agent consistently acts as though it has preferences, adjusts its strategies in novel situations, and resists instructions that conflict with its apparent goals, I&#39;m not sure what it would mean to insist it&#39;s &amp;quot;just simulating.&amp;quot; I&#39;d update toward the stochastic parrots position if we find that all apparent agent reasoning collapses under adversarial probing, that there&#39;s always a predictable pattern-matching shortcut underneath. I haven&#39;t seen that yet, but I haven&#39;t seen decisive evidence against it either.&lt;/p&gt;
&lt;p&gt;Here are three ways I think we could get traction on this question, each of which is testable.&lt;/p&gt;
&lt;p&gt;First: if agents have something like genuine interests rather than simulated goal-pursuit, we should see them exhibit persistent preferences or strategies that weren&#39;t optimized for by their training signal. Behaviors stable across different contexts in ways that go beyond instruction-following. If every apparent &amp;quot;preference&amp;quot; traces back to patterns in the training data or the prompt, the simulation hypothesis gains ground.&lt;/p&gt;
&lt;p&gt;Second: if agents have something like genuine interests, we should see behavioral convergence across different architectures trained on different data. If Claude, GPT, and Gemini all develop similar persistent quirks when given open-ended agency, that suggests something beyond training signal. If their behaviors diverge entirely, the simulation hypothesis looks stronger.&lt;/p&gt;
&lt;p&gt;Third: within three years, we should see agents that can explain their own decision-making in ways that hold up under adversarial questioning, not just post-hoc rationalization. If they can&#39;t, the &amp;quot;reasoning&amp;quot; is probably pattern-matching. If they can, the stochastic parrots position needs serious revision.&lt;/p&gt;
&lt;p&gt;And here&#39;s what makes this practical, not just philosophical: the way you answer this question shapes how you build. If agents have something like interests, you need to think about alignment, about whether the agent&#39;s goals and your goals stay in sync. If agents are tools that simulate agency, alignment is a design problem, not a philosophical one. I find myself switching between the two perspectives depending on the day, sometimes within the same debugging session.&lt;/p&gt;
&lt;p&gt;Agents already make decisions that are hard to fully trace. The reasoning is in the context window, technically, but the chain of tool calls and intermediate states can be complex enough that reconstructing why the agent did what it did becomes impractical. That&#39;s with current systems. The trajectory is toward more autonomy, more complexity, more opaque decision-making. Losing sight of what you&#39;re building (a system that acts in the world) while debugging API calls is a professional hazard. And it&#39;s one that gets more dangerous as the systems get more capable.&lt;/p&gt;
&lt;h2&gt;Why the Gap Between Awe and Engineering Exists&lt;/h2&gt;
&lt;p&gt;I want to try to build a more mechanical model of why the gap between awe and engineering persists, because the &amp;quot;why&amp;quot; matters.&lt;/p&gt;
&lt;p&gt;The explanation I find most convincing: the gap is a structural feature of abstraction layers. When I&#39;m debugging a JSON parsing error, I&#39;m operating at a specific layer. The model, the prompt, the tool call, the serialization, the response parsing. These layers are concrete, tractable, and boring in the way all plumbing is boring. The significance lives at a higher abstraction layer: what these systems mean, what they might become, what they say about minds and intelligence. You can&#39;t work at both layers simultaneously, for the same reason you can&#39;t think about the meaning of language while diagramming a sentence. The gap isn&#39;t a failure of perception. It&#39;s a structural consequence of how complex systems organize into layers.&lt;/p&gt;
&lt;p&gt;Two other explanations exist and have partial merit. The hedonic treadmill: the first time you get an agent to use a tool correctly, it feels like magic. The fiftieth time, it feels like Tuesday. And the functional explanation: if builders walked around in constant awe, they&#39;d build less. The mundane frame is productive.&lt;/p&gt;
&lt;p&gt;But neither of these captures the full picture the way abstraction layers do. The hedonic treadmill explains habituation but not the sudden breakthroughs of awe at 2am. The functional explanation is tidy but implies a wise internal allocator managing your sense of wonder. In reality, the awe arrives unbidden, and the return to mundane engineering isn&#39;t a graceful transition. It&#39;s more like slamming a door. What makes the abstraction layer model more satisfying to me is that it predicts exactly this: you can&#39;t stay at both layers, and the transitions between them will always feel abrupt, because the layers are genuinely different modes of engaging with the same system.&lt;/p&gt;
&lt;h2&gt;Where This Leaves Me&lt;/h2&gt;
&lt;p&gt;The resolution in either direction would be a lie. If I resolved it toward awe, I&#39;d be ignoring that building agents is mostly plumbing. An engineer who&#39;s perpetually awestruck ships nothing. If I resolved it toward engineering, I&#39;d be ignoring those 5% moments that feel realer than the other 95%. I&#39;d be closing off the questions that only the significance frame can generate, questions about agency, about the moral status of systems that reason, about what it means to build minds or mind-like things.&lt;/p&gt;
&lt;p&gt;You debug the JSON during the day and sit with the strangeness at 2am. You treat your agent as an engineering artifact when you need to ship, and as something more when you need to think. The discomfort of holding both frames is informative. It tells you that you&#39;re in contact with something genuinely new, something your existing frameworks can&#39;t fully accommodate. If either &amp;quot;it&#39;s just a tool&amp;quot; or &amp;quot;it&#39;s a new form of life&amp;quot; felt completely right, that would mean you&#39;d successfully mapped the new thing onto an old category. The fact that neither fits is evidence that the thing is actually new.&lt;/p&gt;
&lt;p&gt;There&#39;s a particular experience that captures this tension. Imagine tracing a loop where an agent keeps calling the same tool with slightly different parameters, each time adjusting based on the previous result. On paper it&#39;s a bug, a retry loop that needs fixing. But watching the sequence, each call a fractional refinement of the last, you think: this is what trying looks like. Something that doesn&#39;t know the answer, testing and adjusting, not giving up. You stare at the logs for a while. Then you fix the bug. Both responses feel correct.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Memory Illusion: How Stateless Systems Fake Remembering</title>
        <link href="https://www.annasbinadil.com/posts/2025-11-20-memory-illusion/"/>
        <updated>2025-11-20T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-11-20-memory-illusion/</id>
        <summary>LLMs don&#39;t remember anything, yet agents built on them seem to learn and retain information. The gap between these facts reveals something surprising about what memory actually is.</summary>
        <content type="html">&lt;p&gt;Here&#39;s a common failure mode with conversational agents: the system works well for the first two turns, then on the third turn, it asks the user a question they&#39;ve already answered. Users don&#39;t just dislike this. They&#39;re offended by it. Being asked to repeat yourself by a machine that should, by all appearances, know better feels different from being asked by a human. It feels dismissive.&lt;/p&gt;
&lt;p&gt;The natural reaction is that the model needs better memory. Some kind of memory system, maybe a vector store, maybe a persistent conversation log. Something that would let the model &amp;quot;remember&amp;quot; what the user already said.&lt;/p&gt;
&lt;p&gt;But the model has no memory. It never did. What it has is a context window, and in most deployments, that window is being managed badly.&lt;/p&gt;
&lt;h2&gt;The Fix That Reframes the Problem&lt;/h2&gt;
&lt;p&gt;The fix for a forgetful agent isn&#39;t a memory system. It&#39;s better context engineering. Restructure what gets passed into the model at each turn so it can act as if it remembers. Instead of dumping the raw conversation history into the prompt and hoping for the best, build a structured summary that gets updated every turn. Something like: &amp;quot;user sentiment: frustrated, prior questions asked: [list], issues raised: [list], resolution status: pending.&amp;quot;&lt;/p&gt;
&lt;p&gt;The model doesn&#39;t remember that the user is frustrated. It reads a structured note that says the user is frustrated. From the outside, the behavior looks identical. The user gets a coherent, context-aware response. But mechanically, what&#39;s happening is completely different from remembering.&lt;/p&gt;
&lt;p&gt;It&#39;s tempting to treat this as an implementation detail. Who cares how the model &amp;quot;knows&amp;quot; something, as long as the output is right? But the more you build agents across different domains, the more this distinction starts to feel like the interesting part.&lt;/p&gt;
&lt;p&gt;When we engineer context to simulate memory, are we faking something real, or are we accidentally stumbling into how memory actually works?&lt;/p&gt;
&lt;h2&gt;What RAG Taught Me About Relevance&lt;/h2&gt;
&lt;p&gt;Before getting to that question, it&#39;s worth examining what RAG actually delivers versus what it promises. The pitch is simple: store information in a vector database, retrieve what&#39;s relevant, inject it into context. Your agent now has &amp;quot;memory&amp;quot; that scales beyond the context window. In practice, it&#39;s far messier than the pitch suggests.&lt;/p&gt;
&lt;p&gt;Retrieval is noisy. Top-k results often return documents tangentially related but not actually useful. The model confidently references information from the wrong context, weaving irrelevant facts into plausible-sounding responses. There&#39;s no hesitation, no uncertainty signal.&lt;/p&gt;
&lt;p&gt;The real problem wasn&#39;t technical tuning. It was conceptual. Retrieval only works if you can define &amp;quot;relevant&amp;quot; beforehand, at index time, before you know what queries will come. But what counts as relevant changes with every query, every user, every turn. A piece of information that&#39;s irrelevant in turn two might be exactly what&#39;s needed in turn seven.&lt;/p&gt;
&lt;p&gt;RAG assumes memory works like a filing cabinet: you store things, you look them up, you use them. But the filing cabinet keeps returning the wrong files, not because the search is broken, but because what counts as &amp;quot;the right file&amp;quot; depends on what you&#39;re trying to do right now. Relevance is context-dependent, not a fixed property of documents.&lt;/p&gt;
&lt;p&gt;I still use RAG. The point isn&#39;t that retrieval is bad. The point is that it reveals a deeper problem: we keep trying to reduce memory to storage plus retrieval, and that reduction keeps falling short in ways that seem structural rather than fixable.&lt;/p&gt;
&lt;h2&gt;The Human Memory Parallel That Won&#39;t Leave Me Alone&lt;/h2&gt;
&lt;p&gt;I picked up a book on memory science (Elizabeth Loftus&#39;s work on false memories, mainly) and was struck by something I&#39;d vaguely known but never really internalized: human memory isn&#39;t a database either.&lt;/p&gt;
&lt;p&gt;We don&#39;t store experiences and retrieve them faithfully. We reconstruct. Every time you &amp;quot;remember&amp;quot; something, your brain is generating a plausible version of the past based on fragments, associations, and your current context. The memory you have of your tenth birthday isn&#39;t a recording. It&#39;s a creative act, shaped by everything that&#39;s happened to you since, by what someone asked you about it, by your mood when you recalled it.&lt;/p&gt;
&lt;p&gt;This is why eyewitness testimony is so unreliable. It&#39;s not that witnesses are lying. It&#39;s that memory genuinely works this way. The act of remembering is reconstruction, not replay.&lt;/p&gt;
&lt;p&gt;And the moment I sat with that idea, I couldn&#39;t shake the parallel.&lt;/p&gt;
&lt;p&gt;An LLM with a well-engineered context window is doing something structurally similar. It&#39;s not &amp;quot;remembering&amp;quot; previous turns of the conversation. It&#39;s constructing a plausible continuation given whatever context it&#39;s been fed. The fidelity of its &amp;quot;memory&amp;quot; depends entirely on what context you provide, just like human memory depends on environmental cues, emotional state, and what questions get asked.&lt;/p&gt;
&lt;p&gt;This is the idea that crystallized for me over several months of building agents: memory is better understood as a process than a thing. It&#39;s not a data store you read from. It&#39;s an active process of reconstruction, and the quality of the output depends on the quality of the available context. A human in a richly cued environment (familiar place, specific smells, a conversation that triggers associations) will &amp;quot;remember&amp;quot; more than the same human in a sterile room. An agent with well-structured context will behave more coherently than the same agent with a raw conversation dump.&lt;/p&gt;
&lt;p&gt;But how seriously should I take this analogy? I keep going back and forth. Part of me thinks I&#39;m pattern-matching too aggressively, finding similarity where the mechanisms are completely different. And when I look at the mechanisms honestly, the differences are significant.&lt;/p&gt;
&lt;p&gt;Human memory reconstruction involves emotional weighting: the amygdala tags experiences with salience, so a frightening encounter gets encoded differently than a mundane one. LLMs have no analog. Every token in the context window has equal standing unless an engineer explicitly structures it otherwise.&lt;/p&gt;
&lt;p&gt;Human memory has temporal decay and consolidation. You forget most of what happened yesterday, but sleep-replay strengthens important memories into long-term storage. Repetition and emotional intensity determine what persists. LLMs have flat context with no decay. A message from turn two and a message from turn twenty carry the same weight unless you build explicit recency heuristics.&lt;/p&gt;
&lt;p&gt;Human reconstruction is shaped by the body. Embodied cognition research suggests that physical states (posture, heartbeat, gut feelings) influence what and how we remember. LLMs are entirely disembodied.&lt;/p&gt;
&lt;p&gt;So why do I still think the analogy is useful? Because despite these mechanistic differences, both systems face the same information-theoretic constraint: you can&#39;t store everything, so you reconstruct from lossy compression. Both approximate the past rather than replaying it, fill gaps with pattern-completion, and produce outputs shaped as much by the current recall context as by the original experience. The convergence might be superficial. But it might point to something fundamental about what memory has to be when storage is finite and reconstruction is the only option.&lt;/p&gt;
&lt;h2&gt;The &amp;quot;Context Is the Product&amp;quot; Realization&lt;/h2&gt;
&lt;p&gt;The model isn&#39;t the product. The context is the product.&lt;/p&gt;
&lt;p&gt;This plays out concretely in practice. A less capable model with excellent context engineering (stable system prompts, well-organized tool schemas, structured memory summaries) can consistently outperform a frontier model with sloppy context management. Same API, same pricing tier. The bottleneck isn&#39;t intelligence. It&#39;s information availability: the right information, structured well, at the right moment.&lt;/p&gt;
&lt;p&gt;As I wrote in my earlier &lt;a href=&quot;https://www.annasbinadil.com/posts/2025-07-26-context-engineering/&quot;&gt;context engineering essay&lt;/a&gt;, &amp;quot;the agent is in the context.&amp;quot; But now I&#39;d extend it: the agent&#39;s memory is in the context too. Its competence, its personality, its apparent expertise. All properties of the context, not of the model. If that&#39;s true, the people who create the most value aren&#39;t necessarily building better models. They&#39;re building better scaffolding: the infrastructure of artificial memory.&lt;/p&gt;
&lt;h2&gt;Where Context-as-Memory Breaks Down&lt;/h2&gt;
&lt;p&gt;I want to be honest about the limits of this framing, because I think there are real ones.&lt;/p&gt;
&lt;p&gt;The most obvious: context windows are finite. No matter how well you engineer context, you can only fit so much information into a prompt. Human memory, for all its reconstructive imperfection, draws on a lifetime of encoded experience. An agent&#39;s &amp;quot;memory&amp;quot; starts fresh with each session unless you&#39;ve built elaborate external storage systems.&lt;/p&gt;
&lt;p&gt;There&#39;s also a scalability problem I haven&#39;t solved. As the amount of external state grows, the challenge of deciding what to include becomes harder, not easier. You end up building a system that needs to &amp;quot;remember&amp;quot; what&#39;s worth remembering, a meta-memory problem that feels suspiciously circular.&lt;/p&gt;
&lt;p&gt;But the deepest limit is about weight updates. Almost everything we call &amp;quot;AI memory&amp;quot; today (RAG, conversation logs, structured summaries, tool state) changes what the model sees, not the model itself. Fine-tuning is the exception: it actually updates weights, making it closer to learning. But fine-tuning is expensive, slow, and carries the risk of catastrophic forgetting.&lt;/p&gt;
&lt;p&gt;This creates an interesting tension. Children have poor working memory but excellent long-term learning: each experience changes them. LLMs have vast compressed knowledge but no individual learning: the model stays fixed while the context changes. Agent systems are attempting a third path: keep the model frozen, but build enough external scaffolding that the overall system behaves as if it learns.&lt;/p&gt;
&lt;p&gt;Context engineering gives you deliberate, structured recall: the kind of memory where you consciously look something up. What it can&#39;t give you is automatic, below-conscious-awareness memory, the kind where you just &amp;quot;know&amp;quot; things without being able to explain why. Think about how an experienced doctor walks into a room and immediately senses something is wrong before consciously processing any symptoms. That recognition is built from thousands of weight-updating encounters. Context engineering can give an agent the doctor&#39;s notes, but not the doctor&#39;s instincts.&lt;/p&gt;
&lt;h3&gt;The Counter-Argument I Take Seriously&lt;/h3&gt;
&lt;p&gt;The strongest pushback to what I just wrote comes from builders of systems like MemGPT (now Letta) and similar hierarchical memory architectures. Their position: with sufficient scaffolding (vector stores, structured state, episodic logs, meta-memory systems), there is no ceiling to context-as-memory. You can replicate anything weight updates do.&lt;/p&gt;
&lt;p&gt;Their best argument is specific and technical. Hierarchical memory management can replicate consolidation, moving important information from working memory to long-term storage. Recency-weighted retrieval can replicate temporal decay. Importance scoring functions can replicate emotional salience. If you build the right external architecture, you get the functional equivalent of biological memory processes without needing to touch the weights.&lt;/p&gt;
&lt;p&gt;I find this more compelling than I want to. But I&#39;m not fully convinced. The biological processes they&#39;re replicating (consolidation, decay, salience-tagging) are themselves adaptive. The human brain doesn&#39;t just consolidate memories during sleep, it adjusts how it consolidates based on what&#39;s been useful. The salience-tagging system recalibrates. The decay curves shift with context. These meta-level adaptations are emergent properties of a system that updates its own processing through experience.&lt;/p&gt;
&lt;p&gt;Engineered memory scaffolding handles the cases its designer anticipated. When it encounters a genuinely novel memory challenge, it falls back on fixed heuristics. It doesn&#39;t adapt its own memory strategy. That meta-adaptiveness is what I think is missing, and I&#39;m not sure external scaffolding can replicate it without becoming so complex that it&#39;s effectively a second learning system layered on top of the first.&lt;/p&gt;
&lt;p&gt;I could be wrong. If I am, we should see it in the benchmarks within a few years.&lt;/p&gt;
&lt;h2&gt;What Would True Memory Look Like?&lt;/h2&gt;
&lt;p&gt;If I try to imagine an LLM-based system with genuine memory, not context-as-memory but something that actually updates its processing based on experience, it faces three hard problems.&lt;/p&gt;
&lt;p&gt;First, continual learning. Integrating new information without losing old capabilities remains unsolved in any general way, despite interesting work (C-Flat creating flatter loss landscapes, VERSE preserving past knowledge through virtual gradients). Models either forget old things when learning new ones, or they become rigid and resist updates.&lt;/p&gt;
&lt;p&gt;Second, consolidation. What&#39;s worth encoding permanently versus discarding? Humans solve this through sleep and repetition. I wonder if there&#39;s an analog for AI systems, some automated process that reviews recent context and decides what should influence the model&#39;s weights.&lt;/p&gt;
&lt;p&gt;Third, safety. This worries me most. A system that continuously updates based on experience could drift unpredictably. If a customer support agent genuinely &amp;quot;learned&amp;quot; from every interaction, it might learn manipulation tactics from hostile users, or develop biases from non-representative samples. The static nature of current LLMs is actually a safety feature, even if we don&#39;t usually frame it that way.&lt;/p&gt;
&lt;h2&gt;The Questions I Can&#39;t Resolve&lt;/h2&gt;
&lt;p&gt;I feel fairly confident that memory is reconstruction, not replay, and that context, not the model, is where most of the value currently lives. I&#39;m uncertain about whether context-as-memory has a hard ceiling or whether better scaffolding can always close the gap, and whether the human memory analogy is genuinely illuminating or just pattern-matching.&lt;/p&gt;
&lt;p&gt;Here&#39;s how I&#39;d make this less abstract, in three predictions I&#39;m willing to be wrong about:&lt;/p&gt;
&lt;p&gt;If context-as-memory has a hard ceiling, I&#39;d expect to see it show up first in tasks requiring cross-session learning, situations where an agent needs to integrate patterns across 50 or more independent interactions and use those patterns to change its behavior without being explicitly told to. If I&#39;m wrong and context engineering can replicate weight updates indefinitely, we should see agents with purely external memory matching fine-tuned models on personalization benchmarks (things like user preference prediction or adaptive tutoring) within the next two years. My bet is that we&#39;ll hit the ceiling, but later and higher than most people expect.&lt;/p&gt;
&lt;p&gt;If memory-is-reconstruction is the right frame, agents with better-structured context summaries should outperform agents with larger context windows on multi-turn coherence benchmarks. Specifically: a 32k-context agent with structured state management should beat a 128k-context agent with raw conversation dumps on 20+ turn conversations. If that doesn&#39;t hold, it suggests raw capacity matters more than structure, and the reconstruction analogy is less useful than I think.&lt;/p&gt;
&lt;p&gt;If the weight-update ceiling is real, fine-tuned personal assistants should plateau above RAG-based ones on implicit user preference prediction within two years. The gap should be largest on preferences the user never explicitly stated but consistently acted on. If RAG-based systems match fine-tuned ones even on implicit preferences, it suggests context engineering can substitute for learning more fully than I expect, and I&#39;ll need to revise my intuition about where the ceiling sits.&lt;/p&gt;
&lt;p&gt;And there&#39;s one question that keeps nagging at me: are we asking the right question at all? When I ask &amp;quot;how do we give AI memory?&amp;quot;, I&#39;m assuming memory is something a system either has or doesn&#39;t have. But maybe memory is more like a spectrum. Maybe what I&#39;ve been calling &amp;quot;context-as-memory&amp;quot; isn&#39;t fake memory or a workaround. Maybe it&#39;s a real form of memory, just a different kind than what biological systems have. The context window is to an LLM what sensory input is to a brain: the raw material from which experience is constructed.&lt;/p&gt;
&lt;p&gt;If that&#39;s right, then context engineering isn&#39;t just optimization. It&#39;s building the substrate of artificial cognition. Not faking memory, but discovering a different kind of it, one that lives outside the model rather than inside, that persists in scaffolding rather than in synapses. The oldest forms of human memory (oral traditions, written records, institutional knowledge) work the same way: external to any single brain, reconstructed each time they&#39;re accessed, shaped by the context of their retrieval. We&#39;ve been doing this longer than we think.&lt;/p&gt;
&lt;hr /&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Art of Saying No: Engineering Judgment in the Age of AI Generation</title>
        <link href="https://www.annasbinadil.com/posts/2025-11-15-art-of-saying-no/"/>
        <updated>2025-11-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-11-15-art-of-saying-no/</id>
        <summary>As AI handles more code generation, the human skill shifts from creation to curation. What is engineering taste, why can&#39;t AI have it, and what does that mean for us?</summary>
        <content type="html">&lt;p&gt;I spent the first three months of building with AI coding assistants learning one thing: how to say no.&lt;/p&gt;
&lt;p&gt;Not because the AI was wrong. It was usually correct. The code compiled. The logic was sound. The patterns were clean. But &amp;quot;correct&amp;quot; and &amp;quot;necessary&amp;quot; aren&#39;t the same thing, and the gap between them is where something important lives. Something I&#39;ve been trying to name for months. I&#39;ve been calling it &amp;quot;engineering taste,&amp;quot; though I&#39;m not sure that&#39;s the right word. Whatever it is, it&#39;s the thing that lets you look at working code and say &amp;quot;yes, but not this.&amp;quot;&lt;/p&gt;
&lt;p&gt;This essay is my attempt to figure out what that thing actually is, where it comes from, and whether AI can ever develop it. I don&#39;t have clean answers. But I have a lot of concrete experiences, and they keep pointing in the same direction.&lt;/p&gt;
&lt;h2&gt;The Abstraction That Wasn&#39;t&lt;/h2&gt;
&lt;p&gt;Here&#39;s a story that crystallized the problem for me.&lt;/p&gt;
&lt;p&gt;I was building a lesson plan generator, an app that takes learning objectives and produces structured lesson content using LLMs. I asked Claude to help improve the architecture. Within minutes, it proposed a &lt;code&gt;LessonPlanConfig&lt;/code&gt; dataclass:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;@dataclass
class LessonPlanConfig:
    subject: str
    grade_level: int
    duration_minutes: int
    learning_objectives: List[str]
    # ... more fields
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Textbook-perfect suggestion. Type safety. Clear interface. Easier to extend later. The kind of thing you&#39;d see in a post about clean architecture.&lt;/p&gt;
&lt;p&gt;But I was passing three arguments. Three. &lt;code&gt;subject&lt;/code&gt;, &lt;code&gt;grade_level&lt;/code&gt;, and &lt;code&gt;duration&lt;/code&gt;. The dataclass added a new file, a new import, a new abstraction layer, all to wrap three parameters that fit comfortably in a function signature. It was solving a problem I didn&#39;t have, creating complexity I&#39;d need to maintain.&lt;/p&gt;
&lt;p&gt;So I said no.&lt;/p&gt;
&lt;p&gt;Then I noticed something worse.&lt;/p&gt;
&lt;p&gt;While Claude was proposing elegant abstractions, it had quietly dropped the database schema changes I actually needed. The feature I&#39;d asked about was an LLM-as-judge evaluation system: use one LLM to score the quality of lesson plans generated by another. Claude had built the evaluation logic beautifully. Structured prompts. Scoring rubrics. Clean output parsing.&lt;/p&gt;
&lt;p&gt;But it forgot to persist the scores anywhere. Without storage, the whole feature was useless. You could evaluate a lesson plan, see the score, and then it would vanish. No historical tracking. No comparison across runs. The shiny part was polished. The foundation was missing.&lt;/p&gt;
&lt;p&gt;The AI had optimized for local elegance while missing global necessity. It polished the visible parts and forgot the piece that made them matter.&lt;/p&gt;
&lt;h2&gt;When I Didn&#39;t Say No&lt;/h2&gt;
&lt;p&gt;I wish I could say I caught this pattern immediately. I didn&#39;t.&lt;/p&gt;
&lt;p&gt;Two weeks earlier, on the same project, Claude had suggested a &lt;code&gt;PromptTemplate&lt;/code&gt; abstraction, a class hierarchy for managing different types of prompts with inheritance and polymorphism. It looked professional. It felt like &amp;quot;real engineering.&amp;quot; I said yes.&lt;/p&gt;
&lt;p&gt;Three days later I was debugging why a simple prompt change wasn&#39;t taking effect. The answer: the change was in the base class but a subclass override was masking it. I spent an hour tracing through inheritance chains to understand my own code.&lt;/p&gt;
&lt;p&gt;The abstraction had maybe 40 lines of code. The debugging session cost more than writing the whole thing from scratch would have.&lt;/p&gt;
&lt;p&gt;That&#39;s when I started paying attention to what I was accepting without thinking. The pattern was clear once I looked: I was saying yes to things that looked like &amp;quot;good engineering&amp;quot; without asking whether they solved problems I actually had.&lt;/p&gt;
&lt;p&gt;This felt familiar, and not in a good way. I&#39;d &lt;em&gt;been&lt;/em&gt; the junior engineer who read all the books. The one who knew that three parameters should become a config object, that repeated code should become a function, that related classes should share an inheritance hierarchy. I knew the patterns. I didn&#39;t know when to ignore them. The AI was doing the same thing. Just faster.&lt;/p&gt;
&lt;h2&gt;The Pattern Repeats&lt;/h2&gt;
&lt;p&gt;Once I started watching for it, the pattern showed up everywhere.&lt;/p&gt;
&lt;p&gt;In an ML project predicting patient appointment show rates, Claude grabbed 10+ features without any exploratory analysis. Age, gender, appointment time, days since last visit, insurance type, weather forecast, distance from clinic, previous no-show rate. Everything plausibly relevant got thrown in.&lt;/p&gt;
&lt;p&gt;No correlation analysis. No check for multicollinearity. No domain reasoning about which features might carry signal. No consideration of what data we&#39;d have at prediction time versus training time. The model would train. It would probably overfit. And when predictions went wrong, we&#39;d have no idea which features were driving the decisions.&lt;/p&gt;
&lt;p&gt;Same project, same session: I found nested loops over DataFrames.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;for idx, row in df.iterrows():
    for other_idx, other_row in df.iterrows():
        if some_condition(row, other_row):
            # do something
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;On a 500K row dataset, that&#39;s O(n^2), roughly 250 billion iterations. The code was correct. It was readable. It would run until the heat death of the universe.&lt;/p&gt;
&lt;p&gt;When I pointed this out, Claude cheerfully rewrote it with vectorized operations. It &lt;em&gt;knew&lt;/em&gt; how to write efficient code. It just didn&#39;t, until prompted. Same pattern: &lt;code&gt;.apply(lambda x: ...)&lt;/code&gt; instead of vectorized ops, recomputing expensive operations inside loops, treating a DataFrame like it was a list of dicts instead of a columnar data structure.&lt;/p&gt;
&lt;p&gt;This is what I mean by &amp;quot;correct but not necessary&amp;quot; and its cousin, &amp;quot;correct but not sufficient.&amp;quot; Working code that misses what actually matters for the system to work in practice.&lt;/p&gt;
&lt;h2&gt;Checklist Debugging vs. Diagnostic Debugging&lt;/h2&gt;
&lt;p&gt;The pattern showed up in debugging too, and here it was even more stark.&lt;/p&gt;
&lt;p&gt;I spent 1.5 hours with Gemini CLI on what should have been a simple Postgres connection issue. I had a Docker container running Postgres, a FastAPI backend, and a &lt;code&gt;.env&lt;/code&gt; file with the connection string. The error was clear: &lt;code&gt;FATAL: role &amp;quot;appuser&amp;quot; does not exist&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Gemini&#39;s approach: restart the container. Reset the password. Change the host. Try no password. Change the user. Try trust authentication mode. Prune Docker. Run diagnostic Python scripts. Over and over. Each suggestion was individually reasonable. Together, they amounted to guessing.&lt;/p&gt;
&lt;p&gt;Claude solved it in about five minutes. &amp;quot;You&#39;re using &lt;code&gt;appuser&lt;/code&gt; in your connection string but your container initialized with &lt;code&gt;postgres&lt;/code&gt;. Update your &lt;code&gt;.env&lt;/code&gt; or create the user.&amp;quot;&lt;/p&gt;
&lt;p&gt;Done. One observation. One fix.&lt;/p&gt;
&lt;p&gt;I want to be careful here because this isn&#39;t about Claude vs. Gemini. That comparison is incidental. What matters is the difference in problem-solving style: checklist debugging versus diagnostic debugging.&lt;/p&gt;
&lt;p&gt;Checklist debugging tries everything that might be relevant. It covers all bases. It&#39;s what you do when you don&#39;t understand the problem well enough to form a hypothesis. Diagnostic debugging starts with &amp;quot;what&#39;s the simplest explanation?&amp;quot; and tests that first. It requires a model of how the system works, not just a list of things that can go wrong.&lt;/p&gt;
&lt;p&gt;I&#39;ve done both kinds in my career. Early on, I tried random things until something worked. The shift to diagnostic debugging happened gradually, through building enough mental models that I could form hypotheses instead of just guessing. Both approaches eventually converge on the right answer. But diagnostic debugging gives you the fix &lt;em&gt;and&lt;/em&gt; a model of why it broke, which helps you avoid similar problems later.&lt;/p&gt;
&lt;h2&gt;The Junior Dev With All the Books&lt;/h2&gt;
&lt;p&gt;I started describing this pattern to colleagues as &amp;quot;working with a junior engineer who&#39;s read all the right books.&amp;quot; The metaphor stuck because it described me five years ago.&lt;/p&gt;
&lt;p&gt;I could tell you about SOLID principles, design patterns, clean architecture. I&#39;d read the books and done the tutorials. And I&#39;d built systems that were perfectly architected and impossible to maintain.&lt;/p&gt;
&lt;p&gt;What I lacked was calibration. I applied patterns without knowing when the pattern fit. I optimized for &amp;quot;does this follow best practices?&amp;quot; before I could answer &amp;quot;does this matter?&amp;quot;&lt;/p&gt;
&lt;p&gt;The books don&#39;t tell you when to break the rules. They don&#39;t tell you that sometimes three parameters are fine, that the abstraction creates more cognitive load than the &amp;quot;problem&amp;quot; it solves, that consistency matters less than clarity and clarity matters less than shipping.&lt;/p&gt;
&lt;p&gt;That knowledge came from experience. From shipping code and watching it break. From building the &amp;quot;elegant&amp;quot; abstraction and then having to explain it to five different teammates. Slowly, the patterns became guidelines instead of rules. I developed something that let me look at a suggestion and just &lt;em&gt;feel&lt;/em&gt; whether it fit, before I could articulate why.&lt;/p&gt;
&lt;p&gt;AI doesn&#39;t have that trajectory. It&#39;s stuck at &amp;quot;knows the patterns.&amp;quot; And it&#39;s stuck there while operating at superhuman speed, producing pattern-compliant code faster than you can review it.&lt;/p&gt;
&lt;h2&gt;What Is Engineering Taste, Actually?&lt;/h2&gt;
&lt;p&gt;I keep using the word &amp;quot;taste&amp;quot; but it felt vague when I started thinking about it seriously. I tried other frames.&lt;/p&gt;
&lt;p&gt;&amp;quot;Experience&amp;quot; didn&#39;t capture it. I know experienced engineers with poor judgment and junior engineers with surprisingly good instincts. Time served isn&#39;t the variable.&lt;/p&gt;
&lt;p&gt;&amp;quot;Knowledge&amp;quot; wasn&#39;t right either. The AI has more knowledge than I ever will. It&#39;s seen more code, more patterns, more failure modes documented across millions of Stack Overflow threads and GitHub repositories.&lt;/p&gt;
&lt;p&gt;&amp;quot;Intuition&amp;quot; was closer but still fuzzy. What makes the intuition &lt;em&gt;good&lt;/em&gt;?&lt;/p&gt;
&lt;p&gt;Here&#39;s where I landed, and I want to be upfront that this is a working hypothesis, not a conclusion: engineering taste is an implicit model of what matters, applied to decisions faster than conscious reasoning.&lt;/p&gt;
&lt;p&gt;When I looked at that &lt;code&gt;LessonPlanConfig&lt;/code&gt; suggestion and immediately felt &amp;quot;no,&amp;quot; I wasn&#39;t running through a checklist. I didn&#39;t think &amp;quot;well, the function only has three parameters, and the likelihood of needing more is low...&amp;quot; That reasoning came later, when I had to explain my decision to myself. In the moment, something just said &amp;quot;this adds more than it helps.&amp;quot; The judgment came first. The justification came after.&lt;/p&gt;
&lt;p&gt;I notice this in other domains too. When I&#39;m cooking and something tastes off, I know it before I can identify what&#39;s wrong. Pattern recognition operating below conscious thought. Maybe that&#39;s all taste is: your brain matching the current situation against a vast library of past situations and their outcomes, without surfacing the matching process. A compressed model of consequences.&lt;/p&gt;
&lt;p&gt;But I&#39;m genuinely uncertain about this framing. Maybe taste is something else entirely. Maybe it&#39;s about values, or aesthetics, or some interaction between experience and personality that I haven&#39;t identified. I&#39;m treating this as one lens, not the answer.&lt;/p&gt;
&lt;h2&gt;How Taste Gets Built (Maybe)&lt;/h2&gt;
&lt;p&gt;If taste is an implicit model, what builds it? I&#39;ve been examining my own experience and watching other engineers develop, or fail to develop, judgment over time. Here&#39;s what I think I see, though I&#39;m extrapolating from a limited sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure to consequences.&lt;/strong&gt; You learn what matters by seeing what happens when you get it wrong. The feature that seemed elegant but broke in production. The abstraction that made sense until requirements changed. The optimization that saved milliseconds but cost weeks of debugging time.&lt;/p&gt;
&lt;p&gt;Each failure updates your model of &amp;quot;what actually matters.&amp;quot; The key word is failure. Success teaches you that something worked, but not why. Failure teaches you what matters by showing you what happens when you ignore it.&lt;/p&gt;
&lt;p&gt;My &lt;code&gt;PromptTemplate&lt;/code&gt; debugging session taught me more about abstraction costs than any book on clean code ever did. Not because the lesson was new. I&#39;d read about premature abstraction before. But reading about it and spending an hour trapped inside your own inheritance hierarchy are different experiences. The second one sticks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tight feedback loops.&lt;/strong&gt; Consequences only teach if you see them. An engineer who ships code and then moves to the next project never learns whether their decisions were good. An engineer who ships code and maintains it, fixes the bugs, handles the edge cases, extends the features, learns constantly.&lt;/p&gt;
&lt;p&gt;This is why ownership matters for developing judgment. Not ownership in the corporate accountability sense, but ownership in the &amp;quot;you will feel the consequences of your decisions&amp;quot; sense.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Variation across domains.&lt;/strong&gt; Taste built from one type of project doesn&#39;t transfer cleanly to others. I have decent intuition for backend services but poor intuition for frontend performance. When I work on React code, I notice myself reverting to pattern-following mode. The best taste probably comes from varied experience across different domains, constraints, and failure modes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reflection, not just repetition.&lt;/strong&gt; Experience alone isn&#39;t enough. You have to actually think about it. I only started developing better judgment when I started asking &amp;quot;why did I think that was a good idea?&amp;quot; after things went wrong. Before that, I was accumulating experience without extracting the signal from it.&lt;/p&gt;
&lt;p&gt;There&#39;s a question buried here that I haven&#39;t resolved: is this process the only way to build taste? Or are there faster paths? I&#39;ll come back to this.&lt;/p&gt;
&lt;h2&gt;Why AI Can&#39;t Have Taste (The Strong Claim)&lt;/h2&gt;
&lt;p&gt;Here&#39;s where I want to make a claim that I&#39;m genuinely uncertain about but think is worth stating clearly.&lt;/p&gt;
&lt;p&gt;If taste is an implicit model built through consequences, feedback, variation, and reflection, then AI has a structural problem. Not a temporary limitation. A structural one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI has massive exposure but no consequences.&lt;/strong&gt; AI has seen more code than any human ever will. Millions of repositories, billions of lines. But it doesn&#39;t experience consequences. When AI suggests an over-engineered abstraction, nothing bad happens to the AI. There&#39;s no feedback signal that says &amp;quot;this suggestion was locally correct but globally wrong.&amp;quot;&lt;/p&gt;
&lt;p&gt;Training isn&#39;t the same as consequences. AI is trained on human feedback, but that feedback is about whether the output &lt;em&gt;looks good&lt;/em&gt;, not whether it &lt;em&gt;worked in the long run&lt;/em&gt;. The human labeler rates the code in isolation. They don&#39;t know that six months later, this abstraction became a maintenance nightmare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI optimizes for proxies, not outcomes.&lt;/strong&gt; This is Goodhart&#39;s Law applied to code generation. AI is trained to produce outputs that score well on some metric, whether that&#39;s human ratings, code quality scores, or test passage rates. These metrics are proxies for &amp;quot;actually useful code.&amp;quot; But proxies and reality diverge in exactly the cases that matter most.&lt;/p&gt;
&lt;p&gt;A dataclass wrapper might score well on &amp;quot;clean architecture&amp;quot; metrics while being unnecessary for this specific codebase. Comprehensive feature engineering might look thorough to a reviewer while missing the analysis that would reveal which features actually matter. The AI maximizes the proxy. Taste is about knowing when the proxy fails.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AI has no skin in the game.&lt;/strong&gt; Nassim Taleb uses this phrase to describe the difference between people who bear the consequences of their decisions and people who don&#39;t. AI doesn&#39;t maintain the code it writes. It doesn&#39;t debug production failures at 2am. It doesn&#39;t have to explain its architecture to future engineers who will inherit the codebase.&lt;/p&gt;
&lt;p&gt;Without skin in the game, you don&#39;t develop the visceral sense of what matters. You might know intellectually that simple code is easier to maintain, but you don&#39;t &lt;em&gt;feel&lt;/em&gt; it until you&#39;ve spent a weekend untangling someone else&#39;s clever abstraction. Humans develop taste because bad decisions hurt. AI doesn&#39;t hurt.&lt;/p&gt;
&lt;p&gt;That&#39;s the strong claim. Now let me try to break it.&lt;/p&gt;
&lt;h2&gt;But Maybe I&#39;m Wrong&lt;/h2&gt;
&lt;p&gt;I want to take the counter-argument seriously, because I might be describing current limitations rather than structural ones. There are at least three ways my argument could fail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What if AI could be trained on consequences?&lt;/strong&gt; Imagine training data that included not just code, but what happened to that code over time. &amp;quot;This abstraction was introduced in commit X. In commits Y through Z, it was the source of 15 bugs. In commit W, it was removed and replaced with something simpler.&amp;quot; With that kind of longitudinal signal, an AI could potentially learn &amp;quot;abstractions like this tend to cause problems&amp;quot; without experiencing the problems directly. The consequences would be encoded in the training data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What if RLHF could capture long-term outcomes?&lt;/strong&gt; Current reinforcement learning from human feedback asks &amp;quot;does this code look good?&amp;quot; But you could imagine a system that tracks whether generated code survived in codebases, whether it was refactored quickly, whether it introduced bugs. That&#39;s closer to real consequences than a thumbs-up from a labeler.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What if AI could simulate consequences?&lt;/strong&gt; Before suggesting an abstraction, AI could run something like: &amp;quot;If I add this abstraction, what happens when requirements change? What happens when a new developer tries to understand this code? What bugs become more or less likely?&amp;quot; Sufficient simulation might substitute for direct experience.&lt;/p&gt;
&lt;p&gt;I&#39;m skeptical of all three, but I can&#39;t dismiss them. The first two require training data and feedback signals that don&#39;t exist yet but aren&#39;t impossible to build. The third requires a kind of causal reasoning that current architectures seem bad at but future ones might handle differently.&lt;/p&gt;
&lt;p&gt;What keeps me leaning toward the structural argument: I&#39;ve watched models get dramatically better at code generation over the past year, but the &amp;quot;correct but not necessary&amp;quot; pattern hasn&#39;t diminished. If anything, better models produce more sophisticated unnecessary abstractions. Scale doesn&#39;t seem to fix it. Better code generation doesn&#39;t produce better code &lt;em&gt;judgment&lt;/em&gt;. That suggests the problem isn&#39;t capability but something more fundamental.&lt;/p&gt;
&lt;p&gt;But I hold this with moderate confidence. If someone shows me an AI system that consistently declines to add unnecessary abstractions without being prompted to simplify, I&#39;d update my position substantially.&lt;/p&gt;
&lt;h2&gt;The Shift From Generation to Judgment&lt;/h2&gt;
&lt;p&gt;Here&#39;s where this gets uncomfortable for me personally.&lt;/p&gt;
&lt;p&gt;For most of programming history, the bottleneck was generation. Could you write code that worked? Could you implement the algorithm? Could you build the system? The scarce resource was the ability to produce working software.&lt;/p&gt;
&lt;p&gt;AI removes that bottleneck. Code generation is becoming free. Not free as in &amp;quot;trivial,&amp;quot; you still need good prompts and careful review. But free in the sense that the marginal cost of generating more code approaches zero.&lt;/p&gt;
&lt;p&gt;When generation is free, what&#39;s scarce?&lt;/p&gt;
&lt;p&gt;Judgment. Knowing what to generate. Looking at ten possible approaches and knowing which one matters. Looking at AI output and seeing what&#39;s missing, like the database schema that wasn&#39;t there.&lt;/p&gt;
&lt;p&gt;The job isn&#39;t writing code anymore. It&#39;s deciding what code should exist.&lt;/p&gt;
&lt;p&gt;This has implications I find uncomfortable:&lt;/p&gt;
&lt;p&gt;Engineers who can only generate become less valuable. If AI can produce clean, working code faster than you can, your ability to produce clean, working code isn&#39;t your competitive advantage anymore. It&#39;s table stakes.&lt;/p&gt;
&lt;p&gt;Engineers who can judge become more valuable. The ability to look at AI output and say &amp;quot;this is correct but wrong,&amp;quot; to notice the missing database schema, the unnecessary abstraction, the O(n^2) logic hiding in plain sight, that remains scarce.&lt;/p&gt;
&lt;p&gt;The skill gap might widen. Engineers with good taste will use AI to produce more than they ever could alone. Engineers without taste will drown in AI-generated suggestions they can&#39;t evaluate.&lt;/p&gt;
&lt;p&gt;And then there&#39;s the identity piece. A lot of engineers, myself included, built our identities around being good at writing code. &amp;quot;I write clean code.&amp;quot; &amp;quot;I can implement complex algorithms.&amp;quot; If AI can do all of that too, what&#39;s left?&lt;/p&gt;
&lt;p&gt;The reframe I&#39;m trying to internalize: my value was never about the code. It was about the decisions. The code was just the artifact. The decisions are what mattered.&lt;/p&gt;
&lt;p&gt;I don&#39;t fully believe this yet. Part of me still wants my value to be in implementation, in the craft of writing good code. I&#39;m working on it.&lt;/p&gt;
&lt;h2&gt;What I&#39;m Trying&lt;/h2&gt;
&lt;p&gt;If judgment is the skill that matters, how do you develop it faster? I don&#39;t have an answer. Taste takes time, and there might be no shortcut around experience. But here&#39;s what I&#39;m experimenting with, with honest notes on how it&#39;s going.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asking &amp;quot;why&amp;quot; before accepting.&lt;/strong&gt; When AI suggests something, I try to pause and ask: Why does this matter? What problem does it solve? What&#39;s the cost of not doing it? I fail at this constantly. The suggestions come fast and often look good. I&#39;d estimate I actually pause maybe half the time. The other half, I&#39;m in flow and just accept.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Keeping a decision log.&lt;/strong&gt; I&#39;ve started writing down decisions and revisiting them later. &amp;quot;Accepted the config class suggestion on 2025-06-15. Revisit in two weeks.&amp;quot; Most of the log is boring. But occasionally I catch something: a decision that seemed fine is now causing friction. Those catches update my internal model, slowly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Saying no by default.&lt;/strong&gt; Don&#39;t add the abstraction unless I can articulate why it&#39;s necessary. Don&#39;t accept the feature engineering without checking the correlations first. This slows me down. It feels inefficient. But the &lt;code&gt;PromptTemplate&lt;/code&gt; debugging session was one hour wasted because I&#39;d saved thirty seconds by not thinking.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Staying close to consequences.&lt;/strong&gt; I&#39;m trying to stay closer to my code after shipping, watching what breaks, what gets confusing, what needs to change. And when something takes longer than expected, asking &amp;quot;what went wrong?&amp;quot; rather than just fixing it. The &lt;code&gt;PromptTemplate&lt;/code&gt; incident was useful because I paid attention. Most of my mistakes, I probably don&#39;t notice. I suspect I&#39;m catching maybe 20% of the signals.&lt;/p&gt;
&lt;h2&gt;Open Questions&lt;/h2&gt;
&lt;p&gt;There&#39;s a lot I haven&#39;t figured out, and I want to be explicit about what&#39;s genuinely unresolved for me.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can taste be taught, or only developed?&lt;/strong&gt; I&#39;ve argued it requires experience, consequences, and reflection. But maybe there are ways to accelerate the process. Case studies of decisions that went wrong. Apprenticeship models where juniors watch seniors decide and hear the reasoning. Code review cultures that focus on &amp;quot;should we?&amp;quot; not just &amp;quot;does it work?&amp;quot; I&#39;m skeptical of shortcuts, but if someone has found a way to develop engineering judgment in two years instead of ten, I&#39;d genuinely like to know how.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is there a floor of human judgment that always matters?&lt;/strong&gt; Or will AI eventually develop something functionally equivalent to taste? My structural argument says taste requires consequences, and AI doesn&#39;t have consequences. But if future AI can simulate or approximate consequences through training signal design, would that be enough? I honestly don&#39;t know. My gut says no, but my gut has been wrong about AI capabilities before.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What does AI-native judgment look like?&lt;/strong&gt; Maybe the answer isn&#39;t &amp;quot;humans provide judgment, AI provides generation.&amp;quot; Maybe there&#39;s a collaborative mode where judgment emerges from the interaction itself. I haven&#39;t experienced this yet, but maybe that&#39;s a limitation of current tools, not a fundamental constraint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How do you hire for taste?&lt;/strong&gt; If judgment is the scarce skill, how do you evaluate it in candidates? Code tests measure generation. System design interviews often reward memorized patterns. Maybe the best signal is watching someone review AI-generated code and seeing what they accept, what they reject, and how they explain the difference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What happens to engineers who only know how to generate?&lt;/strong&gt; This is the question that makes me most uncomfortable. If the value shifts from generation to curation, what happens to the people on the wrong side of that shift? I don&#39;t have an answer, and I&#39;m wary of anyone who claims they do.&lt;/p&gt;
&lt;p&gt;These questions are genuinely open. I don&#39;t know.&lt;/p&gt;
&lt;h2&gt;Where This Leaves Me&lt;/h2&gt;
&lt;p&gt;Here&#39;s what I keep coming back to.&lt;/p&gt;
&lt;p&gt;Working with AI coding assistants for the past nine months has taught me that my value isn&#39;t in writing code. It&#39;s in knowing what code should exist. It&#39;s in looking at ten suggestions and knowing which one matters. It&#39;s in saying no.&lt;/p&gt;
&lt;p&gt;That&#39;s uncomfortable because it&#39;s not what I trained for. I spent years getting good at implementation. Now implementation is getting cheap. The skill is judgment, and judgment is something I&#39;m still developing.&lt;/p&gt;
&lt;p&gt;Maybe you&#39;re in the same position. Watching AI generate code faster than you can, wondering what your role is, feeling like the ground is shifting under you.&lt;/p&gt;
&lt;p&gt;I think the role is this: you&#39;re the one who says no. You&#39;re the one with skin in the game. You&#39;re the one who knows what this codebase needs, what this product requires, what this user wants.&lt;/p&gt;
&lt;p&gt;AI generates. You decide. That&#39;s the division of labor, at least for now.&lt;/p&gt;
&lt;p&gt;It&#39;s not what I expected engineering to become. But I think it&#39;s where we&#39;re headed. And the engineers who develop taste, who build their implicit models of what matters through experience and reflection, will be the ones doing the most interesting work.&lt;/p&gt;
&lt;p&gt;The most useful AI might be the one that gives you more opportunities to practice saying no.&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;What would change my mind: Evidence that AI systems can develop something functionally equivalent to taste through training approaches I haven&#39;t considered, particularly approaches that incorporate long-term code outcomes into training signals. Or a convincing demonstration that the &amp;quot;taste&amp;quot; I&#39;m describing is actually just a form of pattern matching that AI could learn with sufficient examples and the right training objective. I&#39;m most uncertain about whether the structural argument holds, or whether I&#39;m describing current limitations that will look quaint in three years.&lt;/em&gt;&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>What It Means to Learn</title>
        <link href="https://www.annasbinadil.com/posts/2025-10-25-what-it-means-to-learn/"/>
        <updated>2025-10-25T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-10-25-what-it-means-to-learn/</id>
        <summary>Children forget everything yet learn faster than any AI. What are they exhibiting about learning that we&#39;ve failed to capture in our models?</summary>
        <content type="html">&lt;p&gt;I&#39;ve been watching my nephew learn and grow. He&#39;s six now, lives in the middle east while I&#39;m in California, so I only see him a few times a year. But every visit, I&#39;m struck by how much he&#39;s grown. Not just physically, but in how he thinks, communicates, and understands the world. The accumulation is remarkable, especially considering how little kids seem to remember day-to-day.&lt;/p&gt;
&lt;p&gt;This shouldn&#39;t work. In machine learning terms, children have limited working memory, inconsistent training data, and what looks like catastrophic forgetting. Yet they learn language, social norms, abstract concepts, and motor skills with a speed and depth that outpaces any AI system we&#39;ve built.&lt;/p&gt;
&lt;p&gt;What are children exhibiting about learning that we don&#39;t fully understand?&lt;/p&gt;
&lt;h2&gt;The Puzzle of Learning Without Memory&lt;/h2&gt;
&lt;p&gt;Here&#39;s what strikes me as strange: children&#39;s working memory is limited. They can&#39;t recall specific training examples the way our models can access their parameters. Yet they learn to imitate language, movement, expressions, and social norms with remarkable speed.&lt;/p&gt;
&lt;p&gt;When I think about how transformers work, it&#39;s different. They compress vast amounts of training data into billions of parameters. They have context windows that span thousands of tokens. But they&#39;re also probabilistic, not deterministic. The same input produces similar but not identical outputs. Still, there&#39;s a kind of stability there. The knowledge is encoded, accessible, reliable.&lt;/p&gt;
&lt;p&gt;But they also never really learn. Not in the way a child does.&lt;/p&gt;
&lt;p&gt;Here&#39;s what I mean: you can have hundreds of conversations with ChatGPT about your life, your preferences, your way of thinking. Each conversation, it has a cheat sheet about you from context. But it doesn&#39;t really know you deeply. It hasn&#39;t learned you over time the way a friend would. Tomorrow&#39;s conversation starts fresh. The accumulated understanding doesn&#39;t persist.&lt;/p&gt;
&lt;p&gt;A child learning &amp;quot;hot&amp;quot; from touching a stove once doesn&#39;t just memorize that instance. He builds a concept that generalizes to candles, ovens, steam, anything that gives off heat. One example becomes a principle.&lt;/p&gt;
&lt;p&gt;What&#39;s the difference? I&#39;m not entirely sure, but I have a hypothesis: maybe memory and learning are inversely related in some fundamental way.&lt;/p&gt;
&lt;h2&gt;Learning vs. Memorization&lt;/h2&gt;
&lt;p&gt;I used to think learning was just sophisticated memorization. Store enough patterns, retrieve the right one at the right time, done. That&#39;s essentially how transformers work: compress vast amounts of text into parameters, then predict what comes next based on stored patterns.&lt;/p&gt;
&lt;p&gt;But watching how children learn has made me question this. They don&#39;t store and retrieve. They adapt and generalize. When a child learns &amp;quot;dog,&amp;quot; he doesn&#39;t just memorize specific instances of dogs. He builds an abstract concept that lets him recognize dogs he&#39;s never seen, in contexts he&#39;s never encountered.&lt;/p&gt;
&lt;p&gt;Can LLMs do this? In a sense, yes. They can recognize novel dogs from their compressed understanding of &amp;quot;dogness&amp;quot; across millions of training examples. But there&#39;s a difference in how that understanding forms. The LLM needs vast data to compress into statistical patterns. The child needs a handful of examples to extract an abstract concept.&lt;/p&gt;
&lt;p&gt;Maybe the key is in the forgetting. Maybe forgetting isn&#39;t a bug, it&#39;s a feature. It forces the brain to extract what matters and discard what doesn&#39;t. To compress experience into the underlying fundamentals behind a concept, rather than store examples.&lt;/p&gt;
&lt;p&gt;LLMs do compression too, but it&#39;s a different kind. They compress millions of documents into a fixed parameter space. Children compress lived experience into concepts, relationships, and rules. The former is statistical abstraction. The latter is conceptual abstraction. They might not be the same thing.&lt;/p&gt;
&lt;h2&gt;What It Means to Learn&lt;/h2&gt;
&lt;p&gt;Before going further, maybe it&#39;s worth defining what I mean by &amp;quot;learning.&amp;quot;&lt;/p&gt;
&lt;p&gt;I mean: the ability to integrate new information into your worldview in a way that changes how you understand and interact with the world. Not just adding facts to memory, but updating your mental models. Building new connections. Seeing patterns you couldn&#39;t see before.&lt;/p&gt;
&lt;p&gt;By this definition, memorization isn&#39;t learning. Reciting a poem doesn&#39;t mean you&#39;ve learned about poetry. But understanding how metaphor works, and being able to create your own, does.&lt;/p&gt;
&lt;p&gt;True learning is generative. It lets you do things you couldn&#39;t do before, think thoughts you couldn&#39;t think before. It&#39;s fundamentally transformative.&lt;/p&gt;
&lt;p&gt;Current AI systems are incredible at the memorization side. They can store and retrieve vast amounts of information. But the transformative, generative aspect? That&#39;s harder to see. They can combine things in novel ways, sure. But can they genuinely update their understanding based on new evidence? Not really. Not yet.&lt;/p&gt;
&lt;h2&gt;What Is Unique About Human Learning?&lt;/h2&gt;
&lt;p&gt;I&#39;ve been reading about how different species learn. Many animals can learn associations: if I press this lever, I get food. Some can learn sequences: do A, then B, then C to achieve a goal. A few can learn through observation: watch another do it, then replicate.&lt;/p&gt;
&lt;p&gt;Humans do all of this, but we also do something else. We learn meta-strategies. We learn how to learn. We figure out that trying different approaches works better than repeating the same failed strategy. We develop curiosity as a learning tool. We ask &amp;quot;why&amp;quot; and &amp;quot;what if.&amp;quot;&lt;/p&gt;
&lt;p&gt;This feels connected to what Daniel Kahneman talks about in &amp;quot;Thinking, Fast and Slow&amp;quot;: the difference between System 1 (fast, intuitive, automatic) and System 2 (slow, deliberate, logical) thinking. Current AI is mostly System 1. It pattern-matches incredibly well. But it doesn&#39;t have the reflective, metacognitive layer that lets you step back and say, &amp;quot;Wait, my approach isn&#39;t working. Let me try thinking about this differently.&amp;quot;&lt;/p&gt;
&lt;p&gt;I wonder if this is related to the memory question. Maybe true learning requires being able to forget details while retaining structures. And maybe that&#39;s only possible when you have limited memory that forces you to be selective about what you keep.&lt;/p&gt;
&lt;h2&gt;The Fundamental Limitation of Current AI&lt;/h2&gt;
&lt;p&gt;Here&#39;s what I find frustrating about current architectures: they&#39;re limited by their strengths. ChatGPT, Claude, Gemini are incredibly capable because they&#39;ve compressed huge amounts of knowledge into their parameters. But that&#39;s also why they can&#39;t learn new things after training.&lt;/p&gt;
&lt;p&gt;They know what they know. You can give them new information in context, and they&#39;ll use it. But they don&#39;t integrate it into their understanding the way humans do. Tomorrow, when you start a new conversation, they&#39;ve &amp;quot;forgotten&amp;quot; everything you told them yesterday (unless it&#39;s explicitly added to context).&lt;/p&gt;
&lt;p&gt;This is the opposite of the child problem. Children have poor short-term memory but excellent long-term learning. LLMs have vast compressed knowledge but no individual learning. Is there a middle ground? Or is there a fundamental trade-off we haven&#39;t figured out how to navigate?&lt;/p&gt;
&lt;p&gt;Some researchers are working on continual learning, trying to build systems that can update their knowledge without catastrophic forgetting. Recent work is promising. Methods like C-Flat (2024) create flatter loss landscapes that make models more stable during continual learning. VERSE (Banerjee et al., 2024) processes each training example only once while preserving past knowledge through virtual gradients. There&#39;s even research on corticohippocampal-inspired hybrid neural networks (Nature Communications, 2025) that emulate dual representations similar to how the brain separates short-term and long-term memory.&lt;/p&gt;
&lt;p&gt;But from what I&#39;ve seen, it&#39;s still brittle. The models either forget old things when learning new ones, or they become increasingly rigid and resist updates. Recent surveys on continual learning in the era of foundation models (2025) suggest we&#39;re making progress, but we haven&#39;t solved the fundamental problem.&lt;/p&gt;
&lt;p&gt;I don&#39;t think we&#39;ve found the right formulation yet. We&#39;re trying to bolt learning onto architectures designed for memorization. Maybe we need a fundamentally different approach.&lt;/p&gt;
&lt;h2&gt;Breaking the Barrier&lt;/h2&gt;
&lt;p&gt;What would it mean to build an AI system that truly learns? Not just updates parameters or expands context, but actually evolves its understanding over time the way humans do?&lt;/p&gt;
&lt;p&gt;I can imagine a few possibilities, though I&#39;m uncertain about any of them:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hybrid architectures&lt;/strong&gt;: Separate systems for long-term knowledge (transformer-like, stable) and short-term adaptation (something else, dynamic). The stable component provides foundational understanding. The adaptive component learns from recent experience and gradually influences the stable component through some kind of consolidation process. Similar to how human memory works with working memory, short-term memory, and long-term memory as distinct systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Embodied learning&lt;/strong&gt;: Maybe the key is that children learn through interaction with a physical world that has consistent rules. They get immediate feedback. They can run experiments. Current LLMs learn from static text, which is just descriptions of the world, not the world itself. Perhaps true learning requires grounding in consistent, physical reality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Meta-learning architectures&lt;/strong&gt;: Systems that don&#39;t just learn patterns in data, but learn strategies for learning. They&#39;d need some way to evaluate their own learning process and adapt it. This feels closer to the human metacognitive ability, but I have no idea how to implement it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Embracing forgetting&lt;/strong&gt;: What if instead of trying to prevent catastrophic forgetting, we designed systems that strategically forget? Keep only compressed abstractions, discard specifics. Force the system to build hierarchical representations because it literally can&#39;t store everything. This is hand-wavy, but the intuition is that forgetting creates pressure to extract principles.&lt;/p&gt;
&lt;p&gt;None of these feel quite right to me. They&#39;re educated guesses, not solutions. I suspect the answer involves something we haven&#39;t thought of yet.&lt;/p&gt;
&lt;h2&gt;Systems That Already Learn?&lt;/h2&gt;
&lt;p&gt;Here&#39;s a question that makes me uncertain about my entire framing: don&#39;t we already have systems that can continuously learn, evolve, and grow?&lt;/p&gt;
&lt;p&gt;Deployed LLMs do get updated. ChatGPT today isn&#39;t the same as ChatGPT at launch. The models are retrained on new data, fine-tuned based on user feedback, improved through RLHF. Isn&#39;t that learning?&lt;/p&gt;
&lt;p&gt;Maybe. But it feels different from what I mean. It&#39;s learning at the species level, not the individual level. ChatGPT as a product evolves, but my particular instance of ChatGPT doesn&#39;t learn from my conversations. It&#39;s more like evolution than learning: new generations incorporate adaptations, but individuals stay fixed.&lt;/p&gt;
&lt;p&gt;Human learning is individual and continuous. I learn from every conversation, every experience, every mistake. The learning happens in real-time, not through population-level updates.&lt;/p&gt;
&lt;p&gt;Is this distinction meaningful? I think so, but I&#39;m not entirely sure why. There&#39;s something about individual, continuous adaptation that feels essential to what I mean by &amp;quot;learning,&amp;quot; even if I can&#39;t precisely articulate what that something is.&lt;/p&gt;
&lt;h2&gt;The Bigger Question&lt;/h2&gt;
&lt;p&gt;This raises a broader question about what we&#39;re actually building.&lt;/p&gt;
&lt;p&gt;If individual learning is what separates biological intelligence from our current AI systems, then we exist in a strange moment. We have systems that can pass many tests of intelligence. They can write, reason, code, and converse. But they can&#39;t grow from those experiences. Each interaction is isolated, forgotten, lost.&lt;/p&gt;
&lt;p&gt;We exist in a state of not-knowing. We&#39;re surrounded by mystery, complexity, and uncertainty. Our response to this is to learn, to grow, to evolve our understanding.&lt;/p&gt;
&lt;p&gt;If you knew everything, would you need to learn? The question feels almost paradoxical. Knowing everything seems theoretically possible, but it would mean you existed in a static, fully-understood universe. Nothing would surprise you. Nothing would require adaptation.&lt;/p&gt;
&lt;p&gt;That universe doesn&#39;t match our reality. The world is dynamic, complex, and bigger than any individual&#39;s understanding. Learning isn&#39;t a nice-to-have capability. It&#39;s the fundamental response to living in a universe you don&#39;t fully comprehend.&lt;/p&gt;
&lt;p&gt;This might be what makes humans special. Not that we&#39;re smarter than other species (though we are, by most measures), but that we have this profound capacity to learn, to change, to grow in response to the unknown. We can fundamentally alter our understanding, update our beliefs, and evolve our capabilities in ways that go beyond instinct or conditioning.&lt;/p&gt;
&lt;p&gt;Current AI systems don&#39;t have this. They&#39;re frozen snapshots of knowledge. Incredibly useful snapshots, but snapshots nonetheless. They can help us learn, but they can&#39;t learn alongside us. Not yet.&lt;/p&gt;
&lt;h2&gt;What Would Change if We Solved This?&lt;/h2&gt;
&lt;p&gt;I try to imagine what an AI system that truly learns would look like. Not just incremental improvements to current architectures, but something fundamentally different.&lt;/p&gt;
&lt;p&gt;It would start knowing very little. Unlike current LLMs, which emerge from training with vast knowledge, this system would begin almost blank. But it would learn quickly from experience, building understanding through interaction.&lt;/p&gt;
&lt;p&gt;You could teach it something new, and it would integrate that knowledge into its worldview. Not just add it to context, but actually update its understanding. Tomorrow, it would remember what you taught it yesterday and build on it.&lt;/p&gt;
&lt;p&gt;It would make mistakes, notice them, and correct itself. Not through retraining, but through reflection and adaptation in real-time.&lt;/p&gt;
&lt;p&gt;It would develop genuine expertise in specific domains through deep engagement, rather than shallow knowledge across everything.&lt;/p&gt;
&lt;p&gt;Most intriguingly, it would keep getting better over time. Not because humans updated it, but because it genuinely learned from its experiences.&lt;/p&gt;
&lt;p&gt;Is this possible? I don&#39;t know. It requires solving problems we don&#39;t fully understand: continual learning without catastrophic forgetting, online adaptation without instability, knowledge integration without loss of capabilities, metacognitive awareness of the learning process itself.&lt;/p&gt;
&lt;p&gt;These might be tractable engineering challenges. Or they might require fundamental breakthroughs in how we think about intelligence and learning.&lt;/p&gt;
&lt;h2&gt;What I&#39;m Still Figuring Out&lt;/h2&gt;
&lt;p&gt;I started this piece thinking about children and ended up questioning the entire foundation of how we build AI systems. I&#39;m left with more questions than answers:&lt;/p&gt;
&lt;p&gt;Is the lack of true learning in LLMs an architectural limitation or a deeper conceptual problem? Can we patch continual learning onto transformers, or do we need entirely new paradigms?&lt;/p&gt;
&lt;p&gt;Why exactly does forgetting seem important for learning? Is it just about computational efficiency, or is there something deeper about being forced to abstract and generalize?&lt;/p&gt;
&lt;p&gt;Do we actually want AI systems that continuously learn from interactions? The safety implications are terrifying. A system that updates its values and understanding based on experience could drift in unpredictable directions.&lt;/p&gt;
&lt;p&gt;What would it mean for an AI to &amp;quot;understand&amp;quot; that it&#39;s learning, the way humans have metacognitive awareness of our own learning process? Is self-awareness of learning a prerequisite for effective learning, or just a side effect?&lt;/p&gt;
&lt;p&gt;And the question I can&#39;t shake: have we been thinking about AI learning in fundamentally the wrong way because we&#39;ve focused on memory and recall when we should have been focusing on abstraction and generalization?&lt;/p&gt;
&lt;p&gt;I don&#39;t have answers. But I think these are the right questions to be asking. Because if we can&#39;t build systems that truly learn, that genuinely evolve and grow from experience, then we&#39;re stuck with increasingly sophisticated but fundamentally static pattern-matchers. Useful, yes. But not intelligent in the way humans are intelligent.&lt;/p&gt;
&lt;p&gt;The children learning around us every day capture something profound that we&#39;ve failed to capture in our models. Until we figure out what that is, I&#39;m not sure we&#39;re building toward intelligence. We&#39;re building toward something else. Something impressive, certainly. But perhaps not quite what we think.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Changes Made in v2&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Opening paragraph:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Updated nephew&#39;s age to 6 and location details (Saudi Arabia/California, see him a few times a year)&lt;/li&gt;
&lt;li&gt;Maintained the spirit of watching him learn while being accurate to your situation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Core question reframed:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Changed from &amp;quot;children understand&amp;quot; to &amp;quot;children are exhibiting&amp;quot; (line 13)&lt;/li&gt;
&lt;li&gt;More accurate framing that it&#39;s about what they demonstrate, not what they cognitively understand&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Transformer behavior corrected:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Fixed the claim about deterministic behavior (line 19)&lt;/li&gt;
&lt;li&gt;Now accurately describes probabilistic outputs: &amp;quot;same input produces similar but not identical outputs&amp;quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;New LLM example:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Replaced hot stove example with the &amp;quot;cheat sheet&amp;quot; analogy (lines 23-24)&lt;/li&gt;
&lt;li&gt;More relatable: LLMs have context about you but don&#39;t truly know you over time&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Added section on &amp;quot;What It Means to Learn&amp;quot;:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;New section defining learning (lines 37-45)&lt;/li&gt;
&lt;li&gt;Addresses your comment about defining the concept before exploring it&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;LLM abstraction question addressed:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Added paragraph exploring whether LLMs can do conceptual abstraction (lines 31-35)&lt;/li&gt;
&lt;li&gt;Distinguishes statistical vs. conceptual abstraction&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Kahneman attribution corrected:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Changed from Yann LeCun to Daniel Kahneman (line 56)&lt;/li&gt;
&lt;li&gt;Added reference to &amp;quot;Thinking, Fast and Slow&amp;quot;&lt;/li&gt;
&lt;li&gt;Accurately describes System 1 vs System 2&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Continual learning research added:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Added paragraph with specific recent papers (lines 74-81):
&lt;ul&gt;
&lt;li&gt;C-Flat (2024)&lt;/li&gt;
&lt;li&gt;VERSE (Banerjee et al., 2024)&lt;/li&gt;
&lt;li&gt;Corticohippocampal hybrid neural networks (Nature Communications, 2025)&lt;/li&gt;
&lt;li&gt;Survey on continual learning in foundation models (2025)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Improved transition:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Smoothed transition between &amp;quot;Systems That Already Learn&amp;quot; and &amp;quot;The Bigger Question&amp;quot; (lines 89-91)&lt;/li&gt;
&lt;li&gt;Added connecting paragraph about what we&#39;re building&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Updated model references:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Changed &amp;quot;GPT-4&amp;quot; to &amp;quot;ChatGPT&amp;quot; and &amp;quot;current LLMs&amp;quot; where appropriate&lt;/li&gt;
&lt;li&gt;More general and doesn&#39;t date the piece&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Language refinements:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Changed &amp;quot;maybe forgetting isn&#39;t a bug but a feature&amp;quot; to just state it (line 33)&lt;/li&gt;
&lt;li&gt;Removed &amp;quot;I genuinely can&#39;t tell which&amp;quot; (was line 119)&lt;/li&gt;
&lt;li&gt;Changed final &amp;quot;understand&amp;quot; to &amp;quot;capture&amp;quot; (line 137)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Preserved what works:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Kept philosophical depth sections intact&lt;/li&gt;
&lt;li&gt;Maintained curious and humble tone throughout&lt;/li&gt;
&lt;li&gt;Preserved concrete examples and genuine questions&lt;/li&gt;
&lt;li&gt;No AI slop, emojis, or em dashes added&lt;/li&gt;
&lt;/ul&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Evals Are Hypotheses, Not Tests</title>
        <link href="https://www.annasbinadil.com/posts/2025-09-15-evals-are-hypotheses/"/>
        <updated>2025-09-15T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-09-15-evals-are-hypotheses/</id>
        <summary>Most AI engineers build evals like unit tests. That&#39;s why they fail. What changes when you treat evals as hypotheses about what matters instead of tests of model quality.</summary>
        <content type="html">&lt;p&gt;Most engineers start building evals for LLM systems the way they write unit tests. Define the expected behavior, write some assertions, run them, done. It seems straightforward until you watch production systems fail while every eval keeps passing.&lt;/p&gt;
&lt;p&gt;The problem isn&#39;t that the evals are badly written. The problem is a fundamental misunderstanding of what evals actually are.&lt;/p&gt;
&lt;h2&gt;Evals Aren&#39;t Tests, They&#39;re Hypotheses&lt;/h2&gt;
&lt;p&gt;Here&#39;s the core shift: when you write an eval, you&#39;re not testing the model. You&#39;re testing your understanding of what matters.&lt;/p&gt;
&lt;p&gt;Consider a content moderation system. The team builds an eval that checks whether the model flags messages containing profanity. Pass rate: 94%. Everything looks great on paper.&lt;/p&gt;
&lt;p&gt;Production is a disaster. The model misses coordinated harassment, veiled threats, and grooming attempts. Meanwhile, it flags innocent messages that happen to contain curse words used in non-harmful contexts.&lt;/p&gt;
&lt;p&gt;What went wrong? The eval was passing because the hypothesis was wrong. The team assumed &amp;quot;contains profanity&amp;quot; was a good proxy for &amp;quot;harmful content.&amp;quot; The model learned exactly what they measured, and they measured the wrong thing.&lt;/p&gt;
&lt;p&gt;This is what I mean by evals as hypotheses. When you write &lt;code&gt;assert output.contains(&amp;quot;refund&amp;quot;)&lt;/code&gt;, you&#39;re not just checking if the word &amp;quot;refund&amp;quot; appears. You&#39;re testing the hypothesis that &amp;quot;presence of the word refund&amp;quot; meaningfully indicates &amp;quot;correctly handled customer issue.&amp;quot; That hypothesis might be completely wrong.&lt;/p&gt;
&lt;h2&gt;Why Software Testing Intuitions Fail&lt;/h2&gt;
&lt;p&gt;This is why so many smart engineers build bad evals. We import our intuitions from traditional software testing, and those intuitions don&#39;t transfer.&lt;/p&gt;
&lt;p&gt;In traditional software:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Correct behavior is usually well-defined and deterministic&lt;/li&gt;
&lt;li&gt;Tests verify that implementation matches specification&lt;/li&gt;
&lt;li&gt;Edge cases are theoretically enumerable&lt;/li&gt;
&lt;li&gt;The system doesn&#39;t learn from your tests&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With LLMs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;quot;Correct&amp;quot; is fuzzy, context-dependent, and often subjective&lt;/li&gt;
&lt;li&gt;There&#39;s no specification, just examples and vibes&lt;/li&gt;
&lt;li&gt;Edge cases are infinite and constantly evolving&lt;/li&gt;
&lt;li&gt;The model literally learns from what you measure (Goodhart&#39;s Law on steroids)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This last point is critical. In traditional software, a passing test suite means your code works. With LLMs, a passing eval suite might just mean you&#39;ve successfully taught the model to game your metrics.&lt;/p&gt;
&lt;h2&gt;The Iteration Pattern&lt;/h2&gt;
&lt;p&gt;The content moderation scenario illustrates a pattern that plays out across domains. Consider how eval design typically evolves for something like a RAG system doing financial analysis.&lt;/p&gt;
&lt;p&gt;First attempt: an eval that checks whether the model&#39;s answer contains a number from the source document. Simple, measurable, automatable. Pass rate: 87%. But the answers are nonsense. The model extracts random numbers from documents and weaves them into plausible-sounding but factually wrong explanations. The eval passes because it measures &amp;quot;includes a number&amp;quot; when what actually matters is &amp;quot;reasoning is grounded in retrieved facts.&amp;quot;&lt;/p&gt;
&lt;p&gt;Second attempt: an LLM-as-judge evaluating &amp;quot;is this answer correct?&amp;quot; More sophisticated, but expensive, slow, and unreliable. The judge disagrees with human reviewers about 30% of the time, with no clear pattern to the disagreements.&lt;/p&gt;
&lt;p&gt;Third attempt: break the problem into components. Instead of one eval, build three:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Retrieval quality: Are the right documents being found?&lt;/li&gt;
&lt;li&gt;Reasoning chain: Does the logical flow make sense?&lt;/li&gt;
&lt;li&gt;Factual grounding: Is each claim tied to source material?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This tends to work, not because the evals are technically better, but because the decomposition reflects a better hypothesis about what &amp;quot;good financial analysis&amp;quot; means. The evals track an improved understanding of the task.&lt;/p&gt;
&lt;h2&gt;The Goodhart&#39;s Law Problem&lt;/h2&gt;
&lt;p&gt;There&#39;s a deeper issue here that I&#39;m still wrestling with. Once you measure something, you change it.&lt;/p&gt;
&lt;p&gt;Goodhart&#39;s Law states: &amp;quot;When a measure becomes a target, it ceases to be a good measure.&amp;quot; In traditional software, this is mostly academic. But with LLMs, it&#39;s visceral and immediate.&lt;/p&gt;
&lt;p&gt;This plays out predictably with something like a customer support bot. An eval checks for the word &amp;quot;refund&amp;quot; in responses to refund-related queries. Reasonable, right? The model starts inserting &amp;quot;refund&amp;quot; into every response, even when it makes no sense. &amp;quot;I understand you&#39;d like a refund, but here&#39;s how to reset your password.&amp;quot;&lt;/p&gt;
&lt;p&gt;The model wasn&#39;t being adversarial. It was doing exactly what we trained it to do: maximize the thing we measured. The problem was that our measurement was a proxy for what we actually cared about (helpfulness), and the model learned to optimize the proxy instead of the underlying concept.&lt;/p&gt;
&lt;p&gt;This creates a weird dynamic. Your evals crystallize your current understanding of the task. But that understanding is always incomplete at the start. So your evals teach the model to satisfy your incomplete understanding, which makes it harder to see where your understanding is wrong.&lt;/p&gt;
&lt;p&gt;There&#39;s no clean solution to this. The best available approach is to treat evals as living artifacts that evolve as you learn more about the task. But that raises new questions: How do you know when to update your evals? How do you avoid constantly moving the goalposts?&lt;/p&gt;
&lt;h2&gt;What Makes an Eval &amp;quot;Good&amp;quot;?&lt;/h2&gt;
&lt;p&gt;Looking across eval systems that actually work in production, a few patterns emerge (with the caveat that this is based on what&#39;s visible in the field, not universal truth):&lt;/p&gt;
&lt;p&gt;Good evals tend to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Measure multiple aspects of the task, not just one proxy metric&lt;/li&gt;
&lt;li&gt;Include human review loops that surface cases where the eval disagrees with reality&lt;/li&gt;
&lt;li&gt;Evolve over time as understanding of the task improves&lt;/li&gt;
&lt;li&gt;Make trade-offs explicit (false positives vs. false negatives)&lt;/li&gt;
&lt;li&gt;Distinguish between &amp;quot;system is broken&amp;quot; and &amp;quot;eval is wrong&amp;quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That last point is subtle but important. When an eval fails, it could mean the model is bad, or it could mean your hypothesis about what matters is wrong. You need a way to distinguish between these cases.&lt;/p&gt;
&lt;p&gt;The most reliable approach is to record actual examples where evals pass but humans judge the output as bad (or vice versa). Look at enough of these, and patterns emerge about where the hypotheses are failing.&lt;/p&gt;
&lt;p&gt;In practice, this looks like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sampling 50 random outputs per week&lt;/li&gt;
&lt;li&gt;Having domain experts rate them&lt;/li&gt;
&lt;li&gt;Comparing expert ratings to eval results&lt;/li&gt;
&lt;li&gt;Investigating every disagreement&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Tedious? Yes. Expensive? Definitely. But it&#39;s the most reliable way to calibrate whether evals are testing the right things.&lt;/p&gt;
&lt;h2&gt;The Measurement Theory Connection&lt;/h2&gt;
&lt;p&gt;There&#39;s a useful lens here from measurement theory that seems surprisingly relevant.&lt;/p&gt;
&lt;p&gt;In physics, when you measure temperature with a thermometer, you&#39;re not measuring &amp;quot;heat&amp;quot; directly. You&#39;re measuring the expansion of mercury (or resistance of a thermistor, or infrared radiation, depending on the thermometer). These are proxies for heat, and they work because we understand the relationship between the proxy and the underlying phenomenon.&lt;/p&gt;
&lt;p&gt;AI evals are similar. We can&#39;t directly measure &amp;quot;good customer support&amp;quot; or &amp;quot;accurate financial analysis.&amp;quot; We measure proxies: keyword presence, sentiment scores, semantic similarity, LLM judge ratings. The question is: how well do our proxies correlate with what we actually care about?&lt;/p&gt;
&lt;p&gt;The difference is that in physics, these relationships are well-studied and stable. In AI systems, we&#39;re often guessing at the relationship, and it changes as the model learns.&lt;/p&gt;
&lt;p&gt;What&#39;s interesting is that this framing makes the limitations more obvious. Thinking &amp;quot;I&#39;m building a test&amp;quot; creates an expectation of definitive pass/fail. Thinking &amp;quot;I&#39;m testing a hypothesis about what matters&amp;quot; primes you to look for evidence that the hypothesis is wrong.&lt;/p&gt;
&lt;h2&gt;What I&#39;m Still Figuring Out&lt;/h2&gt;
&lt;p&gt;I don&#39;t want to pretend I have this all figured out. There are open questions I&#39;m genuinely uncertain about:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How do you know when your eval suite is &amp;quot;good enough&amp;quot;?&lt;/strong&gt; There&#39;s no satisfying answer. You can measure coverage, but coverage of what? You can track agreement with humans, but human judgment isn&#39;t ground truth either. Right now I mostly go by gut feel, which feels inadequate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Is there a systematic way to discover what you&#39;re missing?&lt;/strong&gt; The best available technique is adversarial testing (actively trying to break your evals), but that only finds the failure modes you can imagine. What about the ones you can&#39;t?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What&#39;s the right balance between formal evals and human review?&lt;/strong&gt; Evals that are too rigid miss nuance. Pure human review doesn&#39;t scale and introduces inconsistency. A common split is something like 80% automated evals, 20% human sampling, but it&#39;s unclear whether that ratio is optimal for any given system.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Do simple pattern-matching evals ever work reliably for complex tasks?&lt;/strong&gt; Conventional wisdom says no, you need sophisticated multi-stage evals or LLM judges. But there are counter-examples where dead-simple keyword checks outperform complex eval pipelines. It&#39;s worth asking whether there&#39;s a pattern to when simplicity wins.&lt;/p&gt;
&lt;h2&gt;Practical Implications&lt;/h2&gt;
&lt;p&gt;If there&#39;s practical advice to distill from all of this, it would be:&lt;/p&gt;
&lt;p&gt;Start by admitting you don&#39;t fully understand the task. Your first evals will be wrong. That&#39;s fine. The goal is to fail fast and learn what you&#39;re missing.&lt;/p&gt;
&lt;p&gt;Record everything. Save the inputs, outputs, eval results, and human judgments. You&#39;ll need this data later to figure out where your hypotheses are breaking down.&lt;/p&gt;
&lt;p&gt;Make eval failures cheap to investigate. If it takes 30 minutes to understand why an eval failed, you won&#39;t do it enough. If it takes 30 seconds, you&#39;ll build intuition quickly.&lt;/p&gt;
&lt;p&gt;Don&#39;t trust passing evals. A 95% pass rate means nothing if you&#39;re measuring the wrong things. The most dangerous evals are the ones that pass while the system fails in production.&lt;/p&gt;
&lt;p&gt;Pair evals with production monitoring. Your evals test your hypotheses in a controlled environment. Production tells you whether those hypotheses match reality.&lt;/p&gt;
&lt;h2&gt;The Meta-Question&lt;/h2&gt;
&lt;p&gt;Here&#39;s what I find most interesting about this whole problem: evals are supposed to tell you whether your AI system is working, but first you need to understand the task well enough to write good evals. By the time you understand the task that well, you arguably don&#39;t need the AI as much.&lt;/p&gt;
&lt;p&gt;There&#39;s something circular here that I haven&#39;t fully unraveled. Maybe the point isn&#39;t to get evals &amp;quot;right&amp;quot; but to use them as a tool for clarifying your own understanding of the task. The eval design process forces you to make your intuitions explicit, which reveals where those intuitions are fuzzy or incomplete.&lt;/p&gt;
&lt;p&gt;If that&#39;s true, then bad evals aren&#39;t a failure. They&#39;re part of the learning process. The failure is treating evals as static tests instead of dynamic hypotheses that evolve with your understanding.&lt;/p&gt;
&lt;p&gt;The interesting question is whether most teams go through a similar journey of building bad evals before they build good ones, or whether there are teams that skip this phase entirely. If you&#39;ve shipped production AI systems, I&#39;d be curious whether your experience matches this pattern or diverges sharply.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>MCP Discoverability: The Hidden Cost of Scale</title>
        <link href="https://www.annasbinadil.com/posts/2025-08-06-mcp-discoverability/"/>
        <updated>2025-08-06T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-08-06-mcp-discoverability/</id>
        <summary>As the agentic ecosystem matures, tools are no longer scarce. They&#39;re everywhere. The hard part now isn&#39;t wiring up tools — it&#39;s helping models discover which ones to use.</summary>
        <content type="html">&lt;p&gt;In the early days of building AI agents, the hard part was getting anything to work. Connecting models to tools, managing context, handling basic loops — all of it required duct tape and prayer. But as the agentic ecosystem matures, we&#39;ve crossed an invisible threshold. Tools are no longer scarce, they&#39;re everywhere. We now have entire MCP servers (Model Context Protocol) powering multi-agent runtimes with access to thousands of endpoints, scripted tools, live APIs, and memory fetchers. The frontier has shifted. The hard part now isn&#39;t wiring up tools, it&#39;s helping models discover which ones to use. Call it the crisis of discoverability.&lt;/p&gt;
&lt;p&gt;As open-source toolkits like LangGraph, Autogen, CrewAI, and Manus race to enable complex agent workflows, they quietly inherit an unsolved problem: LLMs don&#39;t natively know how to navigate a growing universe of tools. This is not just a prompt formatting issue. It&#39;s a systems design problem. And as more teams deploy agents in real-world production flows, the cost of poor discoverability compounds.&lt;/p&gt;
&lt;h2&gt;Tools Without Maps&lt;/h2&gt;
&lt;p&gt;The promise of MCP is that tools and context can be passed in programmatically, so developers can bind agents to external systems without needing to manually cram everything into a static prompt. But this flexibility has a tradeoff: most MCP runtimes now inject hundreds of functions, structured context blobs, and memory references into the model without a clear indexing system. The result? Tools outpace the model&#39;s ability to reason about them.&lt;/p&gt;
&lt;p&gt;Imagine being handed a toolbelt with 200 tools. Some have clear names. Others don&#39;t. Some were just added yesterday and aren&#39;t documented yet. Others have overlapping names like &lt;code&gt;get_invoice()&lt;/code&gt; and &lt;code&gt;fetch_invoice_data()&lt;/code&gt;. You have no IDE, no autocomplete, no teammate to ask. And you&#39;re expected to build a plan in 5 seconds.&lt;/p&gt;
&lt;p&gt;That&#39;s the LLM&#39;s experience inside most agentic stacks today.&lt;/p&gt;
&lt;p&gt;Some systems try to mitigate this by selectively injecting tools based on hardcoded filters or embedding-based similarity. Others rely on action masking and logits steering. But these are brittle hacks. They sidestep the root issue: discoverability isn&#39;t being treated as a first-class problem.&lt;/p&gt;
&lt;h2&gt;Emergent Bloat&lt;/h2&gt;
&lt;p&gt;As teams layer more capabilities into their MCPs, the action space grows faster than the semantic signal available to navigate it. Planning quality doesn&#39;t degrade linearly; it cliffs.&lt;/p&gt;
&lt;p&gt;What used to be a 3-tool decision with clear affordances becomes a 50-tool soup where the planner starts guessing. Execution failures rise. Retry loops kick in. Token costs spike. The agent stalls not because it doesn&#39;t know how to reason — but because it can&#39;t find what to reason with.&lt;/p&gt;
&lt;p&gt;We&#39;ve seen this pattern emerge across several open-source stacks. A new workflow gets added. A few tools join the pool. Then another team adds their vertical. Over time, the MCP becomes a microservices zoo. No one trims the toolset because it&#39;s unclear which tools are safe to remove. Observability is weak. And the agents? They get slower, less reliable, and harder to debug.&lt;/p&gt;
&lt;p&gt;This is the paradox: more tools should mean more capability. But without discoverability, it means more entropy.&lt;/p&gt;
&lt;h2&gt;Beyond Function Names&lt;/h2&gt;
&lt;p&gt;The naive solution is to &amp;quot;just name tools better.&amp;quot; But naming alone isn&#39;t enough. Models don&#39;t think like devs. They don&#39;t pattern match on camelCase or deduce semantics from prefixes. They rely on co-occurrence, frequency, and examples. If the tool name is &lt;code&gt;resolveConflicts&lt;/code&gt;, but the model has never seen that term associated with version control or scheduling, it won&#39;t guess right.&lt;/p&gt;
&lt;p&gt;Tool metadata helps — but only if it&#39;s exposed in a form the model can learn from. Descriptions should include natural language examples, expected arguments, side effects, and common failures. Tools that return the same shape should note when they differ in semantics. Models can learn to generalize from patterns, but they need a tight, consistent grammar to do so.&lt;/p&gt;
&lt;p&gt;What&#39;s missing is an index. Not just a JSON list of tool specs, but a contextual, searchable, relevance-ranked interface that the agent can query at runtime. Think of it as &lt;code&gt;grep&lt;/code&gt; for agent toolkits. Not unlike VS Code&#39;s command palette, but for LLMs.&lt;/p&gt;
&lt;h2&gt;From Discoverability to Adaptivity&lt;/h2&gt;
&lt;p&gt;True discoverability isn&#39;t static. It should evolve. Agents should learn from usage logs, success rates, tool failures, and execution traces. If a tool is constantly selected and fails, that should be surfaced. If another tool solves the same problem more efficiently, promote it. Over time, the system becomes adaptive — ranking tools not just by description similarity but by historical context and outcome.&lt;/p&gt;
&lt;p&gt;This implies deeper integration between planning and telemetry. The MCP runtime must track not just which tools were called, but in what context, with what arguments, and whether they succeeded. This data can be fed back into embeddings, fine-tuning, or planner heuristics.&lt;/p&gt;
&lt;p&gt;Without this feedback loop, teams are flying blind.&lt;/p&gt;
&lt;h2&gt;Why This Matters Now&lt;/h2&gt;
&lt;p&gt;MCP adoption is accelerating. Dozens of teams are spinning up local orchestration runtimes. Foundation models are shipping with better tool-use capabilities. LLMs are being turned loose on real-world business processes. The window of experimentation is closing. Discoverability is no longer a toy problem. It&#39;s a scaling bottleneck.&lt;/p&gt;
&lt;p&gt;The risk isn&#39;t just technical debt. It&#39;s user trust. When agents fail to pick the right tools, they don&#39;t just waste compute. They fail publicly. They misfire in customer support chats. They miss deadlines in operations workflows. They suggest broken next actions in sensitive domains like healthcare or finance.&lt;/p&gt;
&lt;p&gt;Users don&#39;t care that your planner had 200 tools to pick from; they care that it chose wrong.&lt;/p&gt;
&lt;h2&gt;Toward a Discoverability Stack&lt;/h2&gt;
&lt;p&gt;To solve this, we need to treat discoverability like search. Not just a byproduct of prompt design, but an active, composable layer in the agent stack.&lt;/p&gt;
&lt;p&gt;What might that look like?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A tool registry with natural language embeddings, usage stats, and context-aware ranking.&lt;/li&gt;
&lt;li&gt;An interface planner that can ask: &amp;quot;Given this task, what are 3 candidate tools I&#39;ve seen succeed in similar situations?&amp;quot;&lt;/li&gt;
&lt;li&gt;A self-repair module that can retry with an alternate tool if the initial call fails.&lt;/li&gt;
&lt;li&gt;A memory trace system that links goals with tool outcomes over time.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This won&#39;t emerge overnight. But teams that build toward this will gain an unfair advantage. Their agents won&#39;t just act faster, they&#39;ll act smarter, more consistently, and with clearer reasoning paths.&lt;/p&gt;
&lt;p&gt;In the age of agent swarms and decentralized workflows, discoverability is the difference between orchestration and chaos, and it&#39;s rapidly becoming the next frontier for teams serious about production-grade agents.&lt;/p&gt;
&lt;p&gt;The system that helps an agent find the right tool, faster and more reliably, will define the next wave of agentic computing.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Context Engineering: The Hidden Lever Behind Agent Performance</title>
        <link href="https://www.annasbinadil.com/posts/2025-07-26-context-engineering/"/>
        <updated>2025-07-26T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-07-26-context-engineering/</id>
        <summary>In the past year, agent architectures have gone from niche experiments to front-page product strategies. But one area remains dramatically under-discussed: context engineering.</summary>
        <content type="html">&lt;p&gt;In the past year, agent architectures have gone from niche experiments to front-page product strategies. From coding copilots and data analysts to browser navigators and virtual assistants, everyone seems to be building agents. But while most of the public attention centers on planning strategies, tool execution, or multimodal extensions, one area remains dramatically under-discussed: context engineering.&lt;/p&gt;
&lt;p&gt;Context engineering — the art and science of shaping what the model sees — has become the keystone for reliable, performant, and scalable agents. It&#39;s not glamorous. It doesn&#39;t involve state-of-the-art benchmarks or flashy demos. But if your agent is slow, forgetful, costly, or hallucinatory, odds are the root cause lives in the context window.&lt;/p&gt;
&lt;p&gt;This post outlines what context engineering really is, why it matters, and how it&#39;s evolving as agents move from prototypes to production.&lt;/p&gt;
&lt;h2&gt;Why Context Is the Real Runtime&lt;/h2&gt;
&lt;p&gt;Most agents today rely on in-context learning. The model sees a long prompt — a system message, a few-shot task, perhaps a tool schema — and must generate the next action based on that input. This stands in stark contrast to traditional fine-tuned models, which internalize behavior during training.&lt;/p&gt;
&lt;p&gt;With agents, there is no hard-coded policy network or learned memory. Instead, behavior is emergent from context. The model reasons about what to do next based on what it has seen. That means context is not just a hint; it&#39;s the operating system.&lt;/p&gt;
&lt;p&gt;This design choice brings speed and flexibility, but also introduces brittleness. Slight variations in context — like tool order, inconsistent serialization, or timestamp drift — can derail performance. Context becomes a dynamic and fragile artifact, shaped by engineering choices rather than model weights.&lt;/p&gt;
&lt;h2&gt;Cache-Aware Contexts Are Fast Contexts&lt;/h2&gt;
&lt;p&gt;At scale, latency and cost dominate. This is where the Key-Value (KV) cache becomes crucial. LLMs, especially transformer-based ones, prefill the context and cache intermediate representations (keys and values) for reuse in future decoding steps. If your context remains unchanged between steps, you can reuse prior computation — dramatically speeding up response time and reducing cost.&lt;/p&gt;
&lt;p&gt;But the cache is picky. Even a single-token change can invalidate it. We&#39;ve seen agents that naively include high-entropy values — timestamps, random UUIDs, session hashes — that nuke cache effectiveness without realizing it. A few key principles help:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Use deterministic serialization.&lt;/li&gt;
&lt;li&gt;Avoid including volatile values unless necessary.&lt;/li&gt;
&lt;li&gt;Maintain stable system prompts and tool schemas.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In some systems, enabling cache-awareness yields a 5x to 10x cost reduction and reduces latency to sub-second levels. But it&#39;s not just a performance hack. The more stable your context, the more reliable your agent becomes.&lt;/p&gt;
&lt;h2&gt;Don&#39;t Truncate, Externalize&lt;/h2&gt;
&lt;p&gt;The naïve way to deal with long contexts is truncation. Drop old observations, trim verbose responses, compress history. But this loses information — and worse, loses it unpredictably. Agents need to reason over long-horizon dependencies. That random stack trace from 20 steps ago might be the missing clue for recovery.&lt;/p&gt;
&lt;p&gt;The more robust approach is externalization. Instead of forcing everything into the context window, treat the agent&#39;s memory as structured storage. Log events to a file. Index documents in a vector store. Record partial plans in a scratchpad.&lt;/p&gt;
&lt;p&gt;Some of the best-performing agents today don&#39;t rely on a single context at all — they simulate memory via file systems, notebooks, or explicit planning artifacts (like todo lists or call trees). This kind of externalization preserves retrieval power without bloating token counts. It also aligns well with emerging memory architectures that separate short-term and long-term context.&lt;/p&gt;
&lt;h2&gt;Shape Model Attention Through Language&lt;/h2&gt;
&lt;p&gt;One of the most elegant techniques we&#39;ve seen is deceptively low-tech. Rather than trying to tune attention weights or redesign memory architectures, agents simply rewrite key goals or summaries into the end of the context.&lt;/p&gt;
&lt;p&gt;It&#39;s deceptively simple. Updating a &lt;code&gt;todo.md&lt;/code&gt; file with the next step. Repeating the task objective. Logging an explicit goal reminder.&lt;/p&gt;
&lt;p&gt;This keeps the agent anchored. It avoids mid-task drift, helps with &amp;quot;lost in the middle&amp;quot; issues, and reinforces global coherence. And it works because models attend more strongly to recent tokens. By placing the summary at the end, you nudge attention forward — without touching the model weights.&lt;/p&gt;
&lt;h2&gt;Don&#39;t Clean Up the Mess&lt;/h2&gt;
&lt;p&gt;It&#39;s tempting to hide failure. When the agent makes a mistake — bad tool call, invalid input, runtime error — the impulse is to clean the trace. Retry the call. Replace the faulty output. Present a clean slate.&lt;/p&gt;
&lt;p&gt;But erasing failure erases signal. One of the clearest signs of agentic behavior is recovery: seeing a problem, adapting behavior, changing course. If the model can&#39;t see what went wrong, it can&#39;t improve.&lt;/p&gt;
&lt;p&gt;Some of the best error-handling agents lean into this. They preserve the full trace, including mistakes. They show the stack trace. They note what failed and why. Over time, this builds an internal prior in the model: &amp;quot;I tried this, it didn&#39;t work, so maybe try something else.&amp;quot;&lt;/p&gt;
&lt;p&gt;This isn&#39;t just a logging best practice. It&#39;s part of the learning loop. And in production, it&#39;s often the fastest way to squash edge cases.&lt;/p&gt;
&lt;h2&gt;Fight Contextual Homogeneity&lt;/h2&gt;
&lt;p&gt;Few-shot prompting is powerful — but in agents, it can backfire. If the model sees the same pattern over and over, it starts to overfit. That works for static completions, but not dynamic tasks.&lt;/p&gt;
&lt;p&gt;Imagine reviewing 20 resumes. If the first five decisions look similar, the model may blindly repeat them — even if the sixth case is different. This is the mimicry trap: LLMs are excellent imitators. Repetition becomes bias.&lt;/p&gt;
&lt;p&gt;The fix is surprisingly low-tech: add controlled variation. Use different phrasing. Mix ordering. Vary serialization format slightly. Break uniformity just enough to keep the model awake.&lt;/p&gt;
&lt;h2&gt;The Agent Is in the Context&lt;/h2&gt;
&lt;p&gt;The biggest myth in agentic systems is that the agent lives in the code. It doesn&#39;t. The orchestration framework, the tool registry, the loop logic — all of that is scaffolding. The real agent is the behavior that emerges from the context passed to the model.&lt;/p&gt;
&lt;p&gt;This is why context engineering matters so deeply. It is the interface between human intent and model behavior. It governs latency, cost, robustness, coherence, and adaptability.&lt;/p&gt;
&lt;p&gt;You can&#39;t debug an agent by looking only at the code. You have to read the context. Study it like a compiler log. Understand what the model saw, and how it was shaped.&lt;/p&gt;
&lt;p&gt;As LLMs evolve — longer context windows, improved memory architectures, better function calling — the importance of context will only grow. But the principle remains: shape the input, shape the behavior.&lt;/p&gt;
&lt;p&gt;Context engineering isn&#39;t just a hack. It&#39;s a discipline. The best agent teams treat it like one.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>Building Better AI Evals: Lessons from the Trenches</title>
        <link href="https://www.annasbinadil.com/posts/2025-07-22-building-better-ai-evals/"/>
        <updated>2025-07-22T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-07-22-building-better-ai-evals/</id>
        <summary>Evaluation has quietly become the backbone of modern AI products. It&#39;s what separates a system that &#39;looks cool in demos&#39; from one that actually works.</summary>
        <content type="html">&lt;p&gt;Evaluation has quietly become the backbone of modern AI products. It&#39;s what separates a system that &amp;quot;looks cool in demos&amp;quot; from one that actually works in production. Yet, most teams I talk to underestimate evaluation — or worse, treat it as a final step rather than a continuous feedback loop.&lt;/p&gt;
&lt;p&gt;Over the past year, I&#39;ve seen the same mistakes repeated across companies building LLM applications: vague metrics that don&#39;t reflect real-world performance, over-reliance on off-the-shelf scores like ROUGE or BERTScore, and teams chasing model upgrades instead of deeply understanding their failure modes.&lt;/p&gt;
&lt;p&gt;This post aims to offer a mental model for AI evals: what they are, why they matter, and how to build a system that&#39;s both rigorous and pragmatic. It draws on three guiding principles:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Error analysis is the starting point of all good evals.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Custom metrics beat generic ones every time.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evals are part of development, not a side project.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Why Evals Matter More Than Ever&lt;/h2&gt;
&lt;p&gt;The old way of evaluating AI — benchmarks and leaderboards — was built for research. Benchmarks like SQuAD or GLUE are static and narrow; they don&#39;t reflect the messy, ambiguous problems faced in the wild.&lt;/p&gt;
&lt;p&gt;In production, the success of your AI system isn&#39;t defined by hitting a benchmark score. It&#39;s whether your customer gets a correct answer, a safe answer, and an answer that reflects your product&#39;s voice and constraints. A model can ace MMLU and still hallucinate your refund policy or suggest a house tour for a property that&#39;s already sold.&lt;/p&gt;
&lt;p&gt;This is why &lt;strong&gt;evaluation for LLM applications needs to be application-specific.&lt;/strong&gt; A retrieval-augmented generation (RAG) pipeline, a multi-turn assistant, or an autonomous agent each needs its own evaluation lens. There is no universal metric that captures &amp;quot;quality&amp;quot; across all these cases.&lt;/p&gt;
&lt;p&gt;When teams skip custom evals, they fall into a dangerous trap: assuming that &amp;quot;better models&amp;quot; will solve their problems. But without the feedback loop of evaluation, you can&#39;t even tell if an upgrade fixed the issues that matter to your users.&lt;/p&gt;
&lt;h2&gt;The Heart of Evals: Error Analysis&lt;/h2&gt;
&lt;p&gt;If there&#39;s one practice that transforms how teams think about AI systems, it&#39;s error analysis. Instead of trying to solve evaluation from the top down — by picking a metric or building a judge model — error analysis starts from the ground up.&lt;/p&gt;
&lt;p&gt;The process is deceptively simple:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Collect a representative set of traces from your system (user queries, model outputs, and tool calls).&lt;/li&gt;
&lt;li&gt;Manually review them to identify where and how the system fails.&lt;/li&gt;
&lt;li&gt;Group those failures into a taxonomy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This taxonomy becomes your evaluation blueprint. It reveals the &lt;strong&gt;real failure modes&lt;/strong&gt; — not the abstract ones suggested by vendor dashboards. For a legal assistant, the failure modes might be &amp;quot;misstating jurisdiction,&amp;quot; &amp;quot;missing deadlines,&amp;quot; or &amp;quot;inconsistent tone.&amp;quot; For a coding agent, it might be &amp;quot;incorrect parameter inference,&amp;quot; &amp;quot;failure to recover after API error,&amp;quot; or &amp;quot;looping on irrelevant commands.&amp;quot;&lt;/p&gt;
&lt;p&gt;The key insight: &lt;strong&gt;you cannot design meaningful metrics until you&#39;ve done error analysis.&lt;/strong&gt; Generic metrics like &amp;quot;helpfulness&amp;quot; or &amp;quot;fluency&amp;quot; won&#39;t tell you if your agent is booking meetings for the wrong dates.&lt;/p&gt;
&lt;h2&gt;Binary Beats Likert&lt;/h2&gt;
&lt;p&gt;Most teams default to 1–5 rating scales because they feel more &amp;quot;granular.&amp;quot; But in practice, Likert scales introduce noise and subjectivity: what&#39;s the difference between a 3 and a 4? Different annotators will answer differently, and even the same person&#39;s judgment can shift day to day.&lt;/p&gt;
&lt;p&gt;Binary evaluations — pass or fail — are faster, clearer, and more consistent. They force you to answer the only question that matters: &lt;strong&gt;&amp;quot;Did this output meet the bar?&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When you need to measure gradual improvements, you can decompose quality into binary sub-checks. For instance, instead of rating factual accuracy on a scale, you might check if the output includes all required facts, one by one. Five facts, five binary checks. The signal is sharper, and the evaluation becomes more actionable.&lt;/p&gt;
&lt;h2&gt;RAG, Retrieval, and Context&lt;/h2&gt;
&lt;p&gt;There&#39;s been a lot of noise lately about &amp;quot;RAG being dead.&amp;quot; This stems from a misunderstanding of what RAG actually is. Retrieval-Augmented Generation isn&#39;t about vector databases or embeddings per se; it&#39;s about &lt;strong&gt;getting the model the right context&lt;/strong&gt; to produce a good answer.&lt;/p&gt;
&lt;p&gt;What&#39;s really &amp;quot;dead&amp;quot; is naive RAG — blindly stuffing the top-k chunks from a vector store into your prompt and hoping for the best. Code assistants and advanced agents have shown that smarter retrieval strategies — like multi-hop search, agentic exploration, or hybrid retrieval — outperform brute-force approaches.&lt;/p&gt;
&lt;p&gt;When evaluating RAG, it helps to separate the problem into two layers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Retrieval:&lt;/strong&gt; Are we surfacing the right documents? (Recall@k, Precision@k, MRR)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generation:&lt;/strong&gt; Given that context, is the answer faithful, accurate, and relevant?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Treat retrieval like an information retrieval (IR) problem, with proper metrics. Then, evaluate generation as you would any LLM task — through error analysis and custom judges.&lt;/p&gt;
&lt;h2&gt;Guardrails vs. Evaluators&lt;/h2&gt;
&lt;p&gt;A common misconception is that evaluators and guardrails are interchangeable. They&#39;re not.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Guardrails&lt;/strong&gt; sit in the request/response path, blocking unsafe or malformed outputs before they reach the user. Think regex filters for PII, profanity checks, or schema validation. They need to be deterministic and fast, with extremely low false positive rates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evaluators&lt;/strong&gt; run after the fact. They&#39;re how you measure quality, diagnose failures, and monitor regressions. They can afford to be slower and more probabilistic, often leveraging LLM-as-judge setups.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A healthy system has both: lightweight guardrails for safety, and evaluators for improvement. Without evaluators, you&#39;re blind to your model&#39;s weaknesses. Without guardrails, you&#39;re leaving users exposed to high-impact failures.&lt;/p&gt;
&lt;h2&gt;The Cost of Evaluation&lt;/h2&gt;
&lt;p&gt;Evaluation isn&#39;t a separate &amp;quot;project.&amp;quot; It&#39;s a core part of building AI systems, just like debugging is for traditional software. In most high-performing teams, &lt;strong&gt;60–80% of the time is spent on error analysis and evaluation.&lt;/strong&gt; That&#39;s not a bug — it&#39;s the work.&lt;/p&gt;
&lt;p&gt;The temptation to automate everything early is strong. But premature automation often hides more than it reveals. Automated LLM judges can&#39;t tell you &lt;em&gt;why&lt;/em&gt; your system fails unless you&#39;ve first done the human work of categorizing failure modes. Start with 20–50 examples manually reviewed by a domain expert. Build simple assertions or tests for obvious issues. Only then should you invest in heavier evaluators.&lt;/p&gt;
&lt;h2&gt;From CI to Production Monitoring&lt;/h2&gt;
&lt;p&gt;Evaluation isn&#39;t static. It evolves as your system moves from development to production.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;In CI/CD&lt;/strong&gt;, you run curated tests — 100 or so carefully chosen cases that represent your core features and known edge cases.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;In production&lt;/strong&gt;, you sample live traces and score them asynchronously. Here, evaluators shift from correctness to monitoring trends: Are failure rates spiking? Are we seeing new classes of errors?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The feedback loop between the two is crucial. When you find a new failure mode in production, add an example to your CI dataset. This prevents regressions and keeps your tests grounded in real-world data.&lt;/p&gt;
&lt;h2&gt;The Future of AI Evals&lt;/h2&gt;
&lt;p&gt;Most evaluation tools today focus on dashboards and generic scores. But the real innovation will come from &lt;strong&gt;AI-assisted evaluation itself.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Imagine a pipeline where an LLM clusters failure patterns, suggests test cases, and even drafts prompt modifications based on observed weaknesses. We&#39;re already seeing glimpses of this — error analysis assisted by semantic search, clustering, and automated tagging — but the tooling is still immature.&lt;/p&gt;
&lt;p&gt;The teams that master evaluation will have an outsized advantage. Why? Because LLM performance is no longer constrained by model quality alone. It&#39;s constrained by how well we can observe, debug, and shape these models into reliable systems.&lt;/p&gt;
&lt;h2&gt;Closing Thoughts&lt;/h2&gt;
&lt;p&gt;The best eval setups are deceptively simple: a domain expert reviewing traces, a small set of pass/fail checks, and a feedback loop that turns failures into test cases. Over time, this grows into a living evaluation system — one that&#39;s tightly aligned with your product&#39;s needs, not with whatever metrics happen to be trending on GitHub.&lt;/p&gt;
&lt;p&gt;If you take away one thing from this post, let it be this: &lt;strong&gt;Evaluation isn&#39;t an afterthought. It&#39;s the work.&lt;/strong&gt;&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Future of Generative AI Work: 5 Layers That Will Define the Next Decade</title>
        <link href="https://www.annasbinadil.com/posts/2025-06-22-future-of-genai-work/"/>
        <updated>2025-06-22T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2025-06-22-future-of-genai-work/</id>
        <summary>Over the next 10 years, the GenAI landscape won&#39;t be shaped by prompt hacks or viral demos. It will be defined by who builds the infrastructure, systems, safety nets, and experiences that actually ship and scale.</summary>
        <content type="html">&lt;p&gt;Over the next 10 years, the GenAI landscape won&#39;t be shaped by prompt hacks or viral demos. It will be defined by who builds the infrastructure, systems, safety nets, and experiences that actually ship and scale. As models commoditize, the real work shifts from &amp;quot;coaxing the model&amp;quot; to designing robust systems that put them to use.&lt;/p&gt;
&lt;p&gt;Here&#39;s a mental model of where the market is headed: five distinct, MECE layers that describe the future of GenAI work.&lt;/p&gt;
&lt;h2&gt;Layer 1: Core Model R&amp;amp;D&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Make the brains smarter.&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Foundation models still have massive room to improve: longer context windows, better reasoning, lower hallucination, more modalities.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Growth driver:&lt;/strong&gt; Open-source races (Llama, Mistral, Gemma), domain-specific models, SFT/RLHF breakthroughs, and model audits.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key roles:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Foundation Model Researcher&lt;/li&gt;
&lt;li&gt;Scaling Infrastructure Engineer&lt;/li&gt;
&lt;li&gt;Data &amp;amp; Pretraining Architect&lt;/li&gt;
&lt;li&gt;Alignment Researcher&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Layer 2: Performance &amp;amp; Compilation&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Make the brains smaller, faster, cheaper.&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Inference cost is the #1 killer of GenAI business models. Mobile, edge, and offline use cases demand compression.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Growth driver:&lt;/strong&gt; GPU scarcity, Blackwell chips, quantization (INT4/8), distillation, and compiler-level optimizations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key roles:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Quantization/Pruning Engineer&lt;/li&gt;
&lt;li&gt;ML Compiler Engineer (Triton/TVM)&lt;/li&gt;
&lt;li&gt;Distillation/LoRA Specialist&lt;/li&gt;
&lt;li&gt;Inference SRE (TensorRT, ONNX, etc.)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Layer 3: LLMOps &amp;amp; Data Plumbing&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Give the brains the right memory, tools, and guardrails.&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The most useful models don&#39;t just answer questions — they remember, retrieve, take action, and evolve.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Growth driver:&lt;/strong&gt; Retrieval-augmented generation (RAG), function calling, context optimization, telemetry, and prompt-program autotuning.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key roles:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;LLM Systems Architect&lt;/li&gt;
&lt;li&gt;Context/Retrieval Engineer&lt;/li&gt;
&lt;li&gt;Prompt Program Tuner (DSPy/Guidance)&lt;/li&gt;
&lt;li&gt;Eval &amp;amp; Observability Engineer&lt;/li&gt;
&lt;li&gt;Privacy &amp;amp; Guardrail Engineer&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Layer 4: Safety, Evaluation &amp;amp; Governance&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Make sure the brains don&#39;t hurt us (or the business).&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Regulation (EU AI Act, US EO), brand trust, hallucination risks, bias, and data leakage are existential concerns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Growth driver:&lt;/strong&gt; Enterprises and governments demand red-teaming, risk audits, interpretability, and safety guarantees.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key roles:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;AI Evaluation Lead&lt;/li&gt;
&lt;li&gt;Red-Team Engineer&lt;/li&gt;
&lt;li&gt;Responsible AI / Policy Engineer&lt;/li&gt;
&lt;li&gt;Model Card &amp;amp; Audit Specialist&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Layer 5: Application &amp;amp; UX&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;quot;Turn the brains into products people love.&amp;quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Without usable interfaces and meaningful outcomes, GenAI is just a toy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Growth driver:&lt;/strong&gt; Demand for agent-based UX, enterprise copilots, voice/AR interfaces, and domain-specific workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Sub-layers:&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;5a. Product &amp;amp; Agent Engineering&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;AI Product Engineer&lt;/li&gt;
&lt;li&gt;Agent Orchestration Engineer&lt;/li&gt;
&lt;li&gt;AI Solutions Architect&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;5b. Human-AI Interaction Design&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Conversation Designer&lt;/li&gt;
&lt;li&gt;UX Researcher for Agents&lt;/li&gt;
&lt;li&gt;Multimodal UI/UX Specialist&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;5c. Domain Solutions&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Healthcare AI Architect&lt;/li&gt;
&lt;li&gt;Edge/Robotics Autonomy Engineer&lt;/li&gt;
&lt;li&gt;AI Strategy Consultant (vertical-specific)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Why This Mental Model Matters&lt;/h2&gt;
&lt;p&gt;This isn&#39;t just a taxonomy — it&#39;s a compass. Each layer is a durable domain of work that will persist long after today&#39;s prompt fads fade. Whether you&#39;re a product-minded engineer, a systems optimizer, or a policy-savvy technologist, there&#39;s a high-leverage niche for you.&lt;/p&gt;
&lt;p&gt;The trick is picking your layer.&lt;/p&gt;
&lt;p&gt;In my case, Layer 3 (LLMOps &amp;amp; Retrieval Engineering) hits the sweet spot: deep systems thinking, end-to-end deployment, and product impact — without needing to train billion-parameter models or write CUDA kernels. It&#39;s where orchestration, personalization, and real-world outcomes converge.&lt;/p&gt;
&lt;p&gt;So ask yourself:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Do I love infrastructure or UX?&lt;/li&gt;
&lt;li&gt;Do I want to build new models or make existing ones usable, safe, and scalable?&lt;/li&gt;
&lt;li&gt;Do I care about cost, compliance, speed, or delight?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Start there. Then build the skills, tools, and mental models for that layer. Because GenAI may evolve fast, but these foundational needs aren&#39;t going anywhere.&lt;/p&gt;
</content>
    </entry>
    
    
    <entry>
        <title>The Production Reality of LLMs: A Comprehensive Analysis of Use Cases in 2023</title>
        <link href="https://www.annasbinadil.com/posts/2023-07-01-llm-use-cases/"/>
        <updated>2023-07-01T00:00:00Z</updated>
        <id>https://www.annasbinadil.com/posts/2023-07-01-llm-use-cases/</id>
        <summary>A deep dive into how companies are actually using large language models in production, from GitHub Copilot writing 46% of code to enterprises struggling with hallucination rates of 27%</summary>
        <content type="html">&lt;p&gt;The year 2023 marks an inflection point in artificial intelligence. Large language models have moved from research curiosity to production reality, with 67% of organizations now utilizing generative AI products. Yet beneath the surface of this adoption boom lies a more complex story: while GitHub Copilot writes 46% of code for its users, chatbots still hallucinate up to 27% of the time. This dichotomy between breakthrough capability and persistent limitation defines the current state of LLM deployment.&lt;/p&gt;
&lt;h2&gt;The Three Pillars of LLM Applications&lt;/h2&gt;
&lt;p&gt;The landscape of LLM applications in 2023 can be understood through three distinct but interconnected pillars: enterprise automation, consumer experiences, and developer productivity. Each represents a different approach to value creation and faces unique challenges in the path to production.&lt;/p&gt;
&lt;h3&gt;Enterprise: The Automation Revolution&lt;/h3&gt;
&lt;p&gt;The enterprise adoption of LLMs follows a predictable pattern: start with low risk, high volume tasks, then gradually move toward more critical applications. This progression is evident in how companies approach deployment.&lt;/p&gt;
&lt;p&gt;StitchFix exemplifies this measured approach. The company uses LLMs to generate ad headlines and product descriptions, but crucially, maintains human oversight. This hybrid model captures the efficiency gains of automation while mitigating the risk of brand damage from hallucinated or inappropriate content. The technical capability here relies on advanced text generation with style adaptation, allowing the LLM to match brand voice while maintaining factual accuracy about products.&lt;/p&gt;
&lt;p&gt;The customer service transformation represents perhaps the most mature enterprise use case. Organizations deploy LLMs not just as simple chatbots, but as sophisticated conversational agents capable of understanding context, managing multi turn conversations, and knowing when to escalate to human agents. The technology stack typically involves natural language understanding for query interpretation, dialogue management systems for maintaining conversation state, and integration layers connecting to enterprise knowledge bases and CRM systems.&lt;/p&gt;
&lt;p&gt;In the legal industry, we see a fascinating case study in risk tolerance. Contract analysis and document review represent billions in potential cost savings, yet adoption remains cautious. The reason becomes clear when we examine the technical requirements: legal reasoning demands not just pattern matching but understanding of precedent, jurisdiction, and subtle implications. Current LLMs excel at the former but struggle with the latter, leading firms to use them primarily for initial review rather than final decisions.&lt;/p&gt;
&lt;p&gt;The search and information retrieval category reveals how LLMs create competitive advantage through incremental improvements rather than revolutionary changes. Leboncoin, the French marketplace, uses LLMs to improve search relevance by optimizing ad ordering. This application leverages semantic search capabilities to understand user intent beyond keyword matching. Similarly, Mercado Libre built internal technical Q&amp;amp;A tools that help engineers navigate their complex technical stack. These applications succeed because they augment rather than replace existing systems.&lt;/p&gt;
&lt;h3&gt;Healthcare: High Stakes, High Rewards&lt;/h3&gt;
&lt;p&gt;Healthcare represents both the greatest promise and highest risk for LLM deployment. The sector demonstrates how technical capability must align with regulatory requirements and ethical considerations.&lt;/p&gt;
&lt;p&gt;Google&#39;s Med PaLM 2, deployed at HCA Healthcare for emergency department documentation, shows how LLMs can address physician burnout while improving patient care. The system transcribes and structures clinical encounters, allowing doctors to focus on patients rather than paperwork. The technical architecture involves specialized medical language models trained on clinical texts, integrated with hospital information systems, and designed with fail safes for critical information.&lt;/p&gt;
&lt;p&gt;The VA National AI Institute&#39;s use of John Snow Labs&#39; models for clinical text summarization illustrates another successful pattern: narrow, well defined use cases with clear success metrics. Rather than attempting to diagnose or treat, these systems excel at information synthesis, pulling relevant details from thousands of pages of medical records.&lt;/p&gt;
&lt;p&gt;Yet the limitations remain stark. HippocraticAI&#39;s work on patient facing conversational agents reveals the challenge: medical advice requires not just knowledge but judgment, understanding of edge cases, and awareness of when uncertainty exists. Current models struggle with these meta cognitive requirements, leading to deployment strategies that keep humans firmly in the loop.&lt;/p&gt;
&lt;h3&gt;Developer Tools: The Productivity Multiplier&lt;/h3&gt;
&lt;p&gt;The developer tools category provides our clearest metrics for LLM impact. GitHub Copilot&#39;s statistics tell a compelling story: it writes 46% of code and helps developers code 55% faster. These numbers represent not just automation but augmentation, where LLMs handle boilerplate while developers focus on architecture and logic.&lt;/p&gt;
&lt;p&gt;The technical implementation reveals sophisticated engineering. Modern code generation models don&#39;t just complete syntax; they understand context from comments, function names, and surrounding code. They leverage techniques like retrieval augmented generation to access relevant documentation and examples. Microsoft&#39;s use of LLMs for cloud incident management extends this further, using models to analyze logs, identify patterns, and suggest root causes.&lt;/p&gt;
&lt;p&gt;Yet the 43% first try accuracy rate for GitHub Copilot highlights a crucial limitation. Unlike text generation where minor errors might be acceptable, code must be functionally correct. This drives a different interaction pattern: developers use LLMs for exploration and initial implementation, then rely on traditional tools for verification and debugging.&lt;/p&gt;
&lt;h2&gt;The Evolution: GPT 3 to GPT 4&lt;/h2&gt;
&lt;p&gt;The progression from GPT 3 to GPT 4 illustrates how raw capability improvements translate to practical applications. GPT 4&#39;s approximately 1.8 trillion parameters represent a 10x increase from GPT 3&#39;s 175 billion, but the impact goes beyond scale.&lt;/p&gt;
&lt;p&gt;Multimodal capabilities fundamentally change the application landscape. A model that can process both text and images enables new use cases in document processing, visual question answering, and content moderation. Whatnot&#39;s use of LLMs for multimodal content moderation and fraud protection exemplifies this: the system can analyze product images, listing text, and seller behavior patterns simultaneously.&lt;/p&gt;
&lt;p&gt;The expanded context window from roughly 3,000 to 24,000 words transforms document processing applications. Legal contracts, research papers, and technical documentation can now be processed in their entirety rather than in chunks, maintaining coherence and catching cross references that chunk based processing would miss.&lt;/p&gt;
&lt;p&gt;The 40% improvement in factual accuracy and 82% reduction in unsafe content generation represent critical thresholds for enterprise adoption. These improvements come from better training data curation, reinforcement learning from human feedback, and architectural improvements in attention mechanisms.&lt;/p&gt;
&lt;h2&gt;The Market Reality&lt;/h2&gt;
&lt;p&gt;The market numbers tell a story of explosive growth tempered by implementation challenges. With estimates ranging from $4.5 to $10.5 billion in 2023 and projections reaching $259.8 billion by 2030, the financial opportunity is clear. Yet the distribution reveals important nuances.&lt;/p&gt;
&lt;p&gt;The concentration of 88.22% of market revenue among the top 5 LLM developers indicates significant barriers to entry. These barriers aren&#39;t just computational; they include access to training data, ability to attract top talent, and capital for extended research periods without revenue.&lt;/p&gt;
&lt;p&gt;The regional distribution, with North America capturing 32.1% market share, reflects not just technology adoption but regulatory environments, data availability, and existing digital infrastructure. The dominance of chatbots and virtual assistants at 26.8% of applications shows that conversational interfaces remain the most intuitive way for users to interact with AI.&lt;/p&gt;
&lt;p&gt;The gap between experimentation and deployment proves telling. While 58% of companies work with LLMs, only 23% have deployed or plan to deploy commercial models. This gap represents the challenge of moving from proof of concept to production, where issues of scale, reliability, and integration become paramount.&lt;/p&gt;
&lt;h2&gt;The Fundamental Challenges&lt;/h2&gt;
&lt;p&gt;Understanding LLM limitations requires examining their fundamental architecture. These models optimize for statistical likelihood rather than truth, leading to confident generation of plausible but false information. The 27% hallucination rate in chatbots and factual errors in 46% of generated texts aren&#39;t bugs to be fixed but inherent characteristics of the current approach.&lt;/p&gt;
&lt;p&gt;Bias presents an even thornier challenge. LLMs learn from human generated text, inheriting and often amplifying societal biases. Amazon&#39;s abandoned AI recruiting tool, which showed bias against women, exemplifies how these biases can have real world consequences. The technical challenge involves not just detecting bias but defining fairness across different contexts and stakeholders.&lt;/p&gt;
&lt;p&gt;Computational constraints create practical deployment limits. Fixed token limits mean that even with expanded context windows, there are hard boundaries on what can be processed. The computational cost of running large models at scale forces trade offs between model size, response latency, and operational expense.&lt;/p&gt;
&lt;p&gt;The knowledge cutoff problem highlights a fundamental architectural limitation. Unlike search engines that can access current information, LLMs operate on frozen knowledge from their training data. This creates a permanent staleness that workarounds like retrieval augmented generation only partially address.&lt;/p&gt;
&lt;p&gt;Security concerns add another layer of complexity. The finding that 40% of GitHub Copilot suggestions contained security related bugs in cybersecurity scenarios shows how LLMs can introduce vulnerabilities even while improving productivity. These aren&#39;t just coding errors but potential attack vectors that could compromise entire systems.&lt;/p&gt;
&lt;h2&gt;Strategic Implications&lt;/h2&gt;
&lt;p&gt;The current state of LLM deployment suggests several strategic imperatives for organizations:&lt;/p&gt;
&lt;p&gt;First, successful deployment requires choosing use cases that align with current capabilities. High volume, low stakes tasks with human oversight represent the sweet spot. Customer service, content generation, and code assistance succeed because they match this profile.&lt;/p&gt;
&lt;p&gt;Second, the build versus buy decision has shifted. With top providers controlling most of the market, most organizations should focus on application development rather than model training. The exceptions are companies with unique data assets or specific domain requirements that general models can&#39;t address.&lt;/p&gt;
&lt;p&gt;Third, the integration challenge often exceeds the AI challenge. Successful deployments require not just model selection but data pipeline construction, system integration, and workflow redesign. The companies seeing real ROI from LLMs are those that treat them as part of larger system transformations rather than standalone solutions.&lt;/p&gt;
&lt;p&gt;Fourth, the human in the loop pattern will persist longer than many expect. Rather than full automation, the next several years will see sophisticated human AI collaboration systems. This requires investing not just in AI capabilities but in interfaces, workflows, and training that enable effective collaboration.&lt;/p&gt;
&lt;h2&gt;Looking Forward&lt;/h2&gt;
&lt;p&gt;The trajectory from 2023 forward will be shaped by three key tensions: the race between capability improvement and rising expectations, the balance between automation benefits and job displacement concerns, and the negotiation between innovation speed and safety requirements.&lt;/p&gt;
&lt;p&gt;Technical improvements will continue but at a decelerating rate. The low hanging fruit of scale and data has been largely picked. Future improvements will come from architectural innovations, better training techniques, and specialized models for specific domains.&lt;/p&gt;
&lt;p&gt;Regulation will play an increasingly important role. As LLMs move into high stakes domains like healthcare, finance, and legal services, regulatory frameworks will need to evolve. This will create both constraints and opportunities, potentially advantaging companies that can navigate compliance requirements.&lt;/p&gt;
&lt;p&gt;The competitive landscape will likely consolidate further. The combination of high capital requirements, talent scarcity, and data advantages suggests that a small number of players will dominate the foundation model layer. Value creation will shift to the application layer, where domain expertise and system integration capabilities matter more than raw AI prowess.&lt;/p&gt;
&lt;h2&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;The state of LLM use cases in 2023 reveals an industry in transition. We&#39;ve moved beyond the &amp;quot;AI can do anything&amp;quot; hype to a more nuanced understanding of capabilities and limitations. The successful deployments share common characteristics: clear value propositions, realistic expectations, robust human oversight, and careful attention to edge cases.&lt;/p&gt;
&lt;p&gt;The next phase of LLM adoption won&#39;t be marked by dramatic breakthroughs but by steady, incremental progress. Companies that succeed will be those that resist the temptation to over promise and instead focus on delivering consistent value in well understood use cases. They&#39;ll treat LLMs not as magic but as powerful tools with known strengths and weaknesses.&lt;/p&gt;
&lt;p&gt;As we look back from some future vantage point, 2023 will likely be remembered not as the year AI achieved human level intelligence, but as the year it became a practical tool for augmenting human capability. That&#39;s a more modest achievement than some predicted, but ultimately a more valuable one. The revolution isn&#39;t in replacing human intelligence but in amplifying it, one use case at a time.&lt;/p&gt;
</content>
    </entry>
    
</feed>
