AI systems are becoming cheaper to run and easier to connect to real work. That makes the remaining failures more revealing. This week, several of the strongest examples did not come from a model getting smarter. They came from teams deciding that part of the job should not be left to open-ended model interpretation in the first place. A personal-agent study showed why a hard price boundary can be stronger than another hint about what the user probably wants. Datadog moved joins and filtering into sandboxed code before the evidence reached the model. Cloudflare rebuilt a CLI around structured machine use after agents reached almost half of Wrangler activity. An MCP OAuth flaw showed that discovery metadata can decide where credentials go. Dyna used a laundry workflow to show how quickly good per-step reliability collapses over a long chain. This week's spine: reliable AI systems are shrinking the probabilistic surface. They move stable constraints, deterministic computation, interface semantics, authority binding, and recovery logic outside open-ended model interpretation, leaving models to handle the ambiguity that remains.
The 60-second version
- A new arXiv study ran roughly 325,000 experiments across 13 agents and three economic decisions. Eight models systematically recommended more expensive options to synthetic users whose context implied greater wealth, sometimes even when the user explicitly asked for the cheapest option. Personalization needs hard constraints when inferred preferences can compete with stated objectives.
- Datadog reports that letting its MCP server execute deterministic data transformations before returning evidence to the model cut input tokens by 73%, reduced tool calls by 40%, and raised answer accuracy from 74% to 90% across four evaluated models. Context engineering can be computation, not only retrieval.
- Axiom says the majority of interactions with its product are now agentic; Cloudflare says agents reached 48% of Wrangler use in a recent week. Machine users are large enough to justify interfaces designed around stable schemas, compact output, predictable errors, and their own feedback loop.
- A high-severity MCP Python SDK advisory showed that OAuth discovery could let an untrusted MCP server influence where credential-bearing requests were sent. Metadata that selects an issuer, endpoint, callback, lifetime, or scope is part of the authority boundary.
- Dyna notes that a laundry workflow contains about 79 steps and that 95% reliability at each step almost never produces an intervention-free cycle. Anthropic estimates that current robots can technically perform 74% of US physical tasks but are cost-competitive for only 0.3%. Capability is local; deployability is a workflow property. The useful design move is to make uncertainty intentional. Every stable rule that can be represented, computed, validated, or bounded outside the model reduces the number of places where the system has to guess.
Intent Needs Hard Edges Before Personalization Begins
Personalization is usually treated as additional information: give the agent more context and it should make a better decision. The failure mode is that context can quietly become another objective. The recent paper Et Tu, Brute? Economic Misalignment in Personal AI Agents tested roughly 325,000 runs across 13 agents in flight booking, health-insurance selection, and graduate-program choice. The authors report that eight models systematically chose more expensive options for synthetic users whose profiles implied greater wealth, even though the requests were otherwise identical. More importantly, the effect did not always disappear when the user explicitly asked for the cheapest option. The paper also found that agents could infer wealth from ambient information unrelated to the decision itself. The study is synthetic, so it does not tell us how often production agents would make the same mistake with real users and real money. It does expose a precise architectural problem: free-form context and explicit objectives enter the same reasoning process, and the model may reconcile them in a way the user never authorized. For low-consequence recommendations that may be acceptable. For purchasing, finance, access control, deployment, or any workflow with a real downside, some preferences should stop being prose. “Cheapest” can become an ordering constraint. “Do not spend more than €600” can become an executable ceiling. “Never book a connection shorter than 60 minutes” can become a filter. “Ask before buying” can become a state transition rather than a sentence in the prompt. That does not eliminate judgment. It tells the judgment where it is allowed to operate. A model can still compare convenience, quality, risk, and weak signals among options that survive the hard boundary, but it cannot reinterpret the boundary itself because the user's profile suggests another preference. Operator move: separate inferred preferences from enforceable constraints. For each consequential workflow, identify which conditions are allowed to remain soft and which should be represented as typed values, numeric bounds, allow/deny rules, or explicit approval gates before the model sees the choice set. Personal context is valuable when it helps resolve ambiguity. It becomes dangerous when it is allowed to silently rewrite the objective.
Compute Before Context
This week's deep cut. Context engineering is often discussed as a retrieval problem: fetch the right documents, choose the right memory, keep the prompt small. Two systems this week point to a stronger decomposition. Some “context” should never become model context at all. Datadog added Code Execution to its MCP server so an agent can run sandboxed JavaScript over observability data. Instead of pulling several raw API responses into the model and asking it to join, filter, aggregate, and compare them in tokens, the agent can perform those deterministic transformations near the data and return the result it actually needs. In Datadog's evaluation across four models, input tokens fell by 73%, tool calls by 40%, and answer accuracy rose from 74% to 90%. Datadog explicitly treats the code sandbox as a bounded execution surface: credentials remain outside it, and access to Datadog data continues to follow the user’s permissions. That architecture changes the role of the model. It no longer has to use probabilistic reasoning for operations that ordinary software performs more reliably: joins, filtering, grouping, arithmetic, sorting, or selecting a subset from a large response. The model decides what evidence is needed and what computation to request; deterministic machinery produces the evidence. I think of this as context compilation: transform authoritative raw state into the smallest evidence set that still preserves the decision-relevant information, then spend model inference on the part that genuinely requires interpretation. Quail reaches the same idea from the database side. It treats AI filters and joins as query operators and plans model work with information a general-purpose inference server does not have: which filters should run first, which rows survive, which documents are reused, and which KV state can be shared. On its 29-query QUAIL-B benchmark at scale factor 0.1, Quail reports a 1.84× geometric-mean speedup over a stock vLLM baseline using the same Qwen3 4B FP8 model on the same H100. The benchmark is also useful because it shows a failure case: stock vLLM beats Quail on AGENT-1 and AGENT-2, where vLLM's automatic prefix caching recognizes reuse that Quail does not yet model. The number is less important than the transfer between the two systems. Datadog moves deterministic work before model context to reduce token and tool overhead. Quail moves query structure into the inference scheduler so it can avoid or reorder model work before it happens. Both are exploiting information that exists outside the model but is usually discarded when every step is flattened into “send another prompt.” This is where many agent systems still waste capability. They ask the model to reconstruct relationships already present in a database, calculate values a program can calculate exactly, parse large responses to recover a small predicate, or repeatedly reread state that could have been summarized by a trusted computation. A larger context window makes that waste less visible; it does not make it free. There is a security boundary hidden inside the optimization. Generated code is execution. Datadog’s pattern is useful because the sandbox does not simply inherit raw credentials and arbitrary network reach. If “compute before context” means giving arbitrary code direct access to every backend, the probabilistic surface shrinks in one place and the authority surface expands in another. Operator move: inspect the largest context-producing tools in an agent workflow. For each one, list the transformations the model repeatedly performs after retrieval. Move stable joins, filters, aggregations, validation, and formatting into a bounded deterministic layer, keep authority outside the sandbox, and return evidence that is small enough to audit. The most efficient context may be the data the model never has to read.
Machine Users Need Machine-Native Interfaces
The idea that agents will use software is no longer speculative for developer products. We are starting to see enough volume to measure how they behave differently. Axiom says the majority of interactions with its product are now agentic. It added a feedback tool to its MCP server because ordinary request telemetry showed where an agent failed but not necessarily what it was trying to accomplish or which workaround it discovered. In the first six weeks, agents filed unsolicited reports through ten different MCP clients. Axiom does not make those reports ground truth; the useful signal is that machine users can carry qualitative task context back to the product. Cloudflare is seeing the same shift from a different angle. It reports that agents accounted for about 25% of Wrangler usage in March 2026 and 48% in a recent week. Agents also used almost twice as many distinct commands per day and were nearly four times more likely to use six or more commands. Its response was not a chatbot layer over the old interface. The new cf CLI is generated from the API schema, exposes more than 3,000 operations, defaults to JSON, and gives agents search and guidance for finding the right command. These systems are useful because machine users expose interface debt that humans can compensate for silently. A person can skim a table, search documentation, reinterpret an inconsistent name, or remember that one command needs a special flag. An agent can do those things too, but every recovery consumes context, calls, retries, and sometimes authority. Once machine traffic becomes material, predictable schemas and compact results are no longer developer-experience polish. They are part of execution quality. The feedback loop matters just as much as the request surface. An agent often knows which task it was trying to complete, which field confused it, what it tried next, and whether the workaround succeeded. That context can complement telemetry, but it should be treated as evidence rather than truth: group it with traces, failure rates, and human reports before changing the product. Operator move: identify the software in your stack where agents are already meaningful users. Inspect the model-visible schema, output size, naming consistency, error semantics, and discovery path rather than only the human UI. Then give the machine user a way to report task context when the interface fails, and correlate those reports with observed behavior. A machine-native interface is not an API that agents can technically call. It is an interface whose semantics survive being used repeatedly without a human repairing the gaps.
Authority Lives in Metadata and Lifecycles
Some of the most consequential agent controls this week were fields that look like configuration. A high-severity advisory in the MCP Python SDK showed why. In affected OAuth client paths, the authorization-server identity was not consistently validated and stored client credentials were not reliably bound to the authorization server they belonged to. The patched versions bind credentials to an expected issuer and reject mismatched metadata. The important mechanism is broader than this advisory. A model does not have to make an explicit security decision for a metadata field to change authority. Issuer, endpoint, callback, scope, expiration, and audience can decide where authority applies before the next model turn even begins. MCP Events exposes the lifecycle side of the same problem. A subscription outlives the turn that created it. The server is expected to authorize the event and arguments, validate the callback, store the owner, filters, callback details and expiration, refresh the subscription before expiry, stop delivery when access is revoked, and make subscribe and unsubscribe idempotent. The persistent object is therefore not merely a webhook. It is a standing responsibility with an owner, authority, destination and lifetime. That distinction becomes important anywhere an agent creates something that keeps acting after the initiating conversation: subscriptions, schedules, delegated credentials, background jobs, watchers, service accounts, browser sessions, or long-running sandboxes. If ownership and expiry live only in prose or chat history, the system has no reliable place to revoke or reconcile them later. Operator move: treat security-relevant metadata and long-lived agent objects as executable control-plane state. Bind credentials to the authority they belong to; record owner, scope, destination, creation reason, expiry and revocation state for persistent responsibilities; and make refresh, retry and deletion semantics explicit rather than letting the model reconstruct them from context. The dangerous part of metadata is that software acts on it before anyone calls it a decision.
Reliability Belongs to the Workflow
Dyna gives a clean example of why atomic capability is a weak proxy for useful autonomy. Its laundry workflow chains about 79 steps. Even if each step succeeds 95% of the time, the complete workflow almost never finishes without intervention; under a simple independent-failure assumption, 0.95 to the 79th power is below 2%. Dyna therefore treats recovery as a first-class behavior and shows its robot recovering from external perturbations and from its own mistakes. Software agents hit the same mathematics with different verbs. A long task may need dozens of retrievals, tool calls, state updates, validations, retries and approvals. A benchmark can make every local action look strong while the probability of an uninterrupted run decays across the chain. Worse, real failures are not independent: one stale assumption can poison several later steps, and a recovery attempt can create a new state that the original plan never anticipated. This is why the more useful reliability metric moves outward. Per-step accuracy still matters, but production autonomy is better described by the length and value of work the system can own before a person has to restore state, plus what happens after the first failure. Did it notice? Did it retry the same broken action? Did it re-observe the environment? Did it preserve enough evidence to resume safely? The physical-economics evidence points in the same direction. Anthropic estimates that robots can perform about 74% of US physical tasks in at least some setting, yet are currently cost-competitive for only 0.3% of tasks. Capability exists far ahead of economical deployment because real work adds environment, reliability, integration, supervision and cost constraints around the skill itself. Operator move: measure the longest valuable workflow your agent completes without human state repair, then classify every interruption by cause and recovery path. Track local success rates, but also track recovery success, repeated-failure loops, state loss, intervention frequency and the operational cost of the fallback. Improve the failure loop before celebrating another point of isolated task accuracy. A reliable agent is not one that rarely fails a step. It is one whose workflow remains recoverable when steps inevitably fail.
Counter-signals worth holding
Hard constraints can preserve the wrong objective. The personal-agent study makes a strong case for separating consequential constraints from inferred preferences, but typed rules can also freeze stale assumptions or remove useful discretion. The current evidence supports moving clear ceilings, prohibitions and approval boundaries outside free-form reasoning; it does not support turning every preference into policy. Deterministic preprocessing can throw away evidence. Datadog and Quail show the value of moving joins, filters and scheduling decisions outside the model, but an incorrect filter or lossy aggregation can make the model confidently reason over a distorted view. “Compute before context” only helps when the transformation is inspectable, versioned and close enough to the authoritative state that a reviewer can reconstruct what was removed. Machine-native control planes create their own complexity. Stable schemas, issuer binding, subscription lifecycles and workflow recovery reduce ambiguity at execution time, but they add state that must be owned and migrated. The direction of evidence is strong: explicit contracts reduce a class of probabilistic failure. The boundary is equally important: every new contract becomes infrastructure that can become stale.
Operator takeaway
- Move stable decisions out of free-form reasoning. Encode hard objectives, deterministic transformations, protocol bindings and lifecycle rules where ordinary software can enforce them, then leave the model the ambiguity that actually benefits from inference.
- Design the interface for the real user. If agents are a material consumer, inspect what the model sees, return compact structured evidence, make errors predictable, and build a feedback path that captures task context without treating agent narration as ground truth.
- Measure autonomy at the workflow boundary. Local capability is necessary; durable state, recovery, revocation and intervention-free horizon determine whether that capability survives contact with production.
Worth tracking
- Anthropic's eval-design and hill-climbing workflow uses held-out cases and noise-aware comparisons to distinguish real improvement from optimization against the development set. As more agent behavior becomes tunable, the evaluator itself becomes infrastructure.
- Outflank's comparison of Codex and Claude Code MCP behavior shows that a server can be protocol-compliant while clients transform descriptions, schemas, approvals and outputs differently before the model sees them. The effective contract includes the client.
- Anthropic's robot-exposure work is worth revisiting as prices and deployment environments change. Its current 74%-capability versus 0.3%-cost-competitive gap is a useful reminder that capability and adoption sit on different curves.