NoteEngineering leadership

AI Scaling Is Becoming a Selection Problem

As autonomous execution gets cheaper, the bottleneck shifts from raw capacity to selection: which work deserves to run, how hard it should press, what evidence can reject it, and where specialization fits. This issue shows how backpressure, independent acceptance, effect-based authority, and feedback-aware model routing turn abundant AI execution into controlled production systems.

Lukman Nuriakhmetov
Lukman Nuriakhmetov
15 min read · September 28, 2026

AI systems are getting more capable, but this week the more useful pattern was about restraint. One set of teams was adding more parallel agents. Databricks was looking for customer failures that ordinary service health missed. Datadog moved repeated investigative work from a frontier teacher into a specialized 9B student. A DFDX Labs experiment produced an extraordinary amount of software and almost no demand. The common problem is no longer simply how to get an AI system to do more work. Once execution becomes cheap enough, the system needs to become selective about which work deserves to start, how much pressure to apply, what evidence can reject the result, where authority must stop, and which model should execute the recurring loop. This week's spine: as autonomous execution gets cheaper, scaling becomes a selection problem. Reliable systems need admission rules, backpressure, independent acceptance evidence, effect-based authority boundaries, and execution paths specialized to the feedback available in each task.

The 60-second version

  • Toby Ord's analysis of published swarm results suggests that parallel agents can buy lower wall-clock time at a compute premium: a four-agent swarm needed roughly twice the total reasoning tokens for the same performance while potentially finishing in about half the time. Parallelism is an economic choice, not a default maturity step.
  • Canva's queue workers used local success and failure feedback to reduce concurrency during a 32.5-hour overload. The protected fleet kept completing about two million messages per hour; the incident recorded about 1.8 million failed attempts while the dead-letter queue grew by only 22 messages. Autonomous capacity also needs a brake.
  • Databricks' RADAR catches gray failures from user-error behavior even when ordinary component health remains green. NVIDIA's agent-evaluation guidance makes the same distinction from another angle: verify consequences in the environment, not only the agent's trajectory or final message.
  • Datadog trained a Qwen3.5-9B student on successful investigation traces from GLM-5.3. On held-out internal incidents it reached 0.55 Recall@5 versus 0.63 for the teacher, at about $0.003 per investigation versus $0.06. Some recurring work is becoming a training and serving problem rather than a permanent frontier-inference problem.
  • DFDX Labs let an agent operate a business for three weeks. It built 17 paid products and 46 API endpoints, consumed more than five billion tokens, and generated $1.54 in revenue. Cheap supply does not manufacture authorized demand. The practical shift: stop treating every unit of available AI capacity as work that should be consumed. Design explicit mechanisms for admitting, throttling, rejecting, bounding, and specializing execution.

Parallelism Needs a Brake

Multi-agent systems are often presented as an obvious next step: if one capable agent helps, several working in parallel should help more. The missing variable is the price of elapsed time. Toby Ord reconstructed swarm-scaling behavior from published OpenAI results and estimated task-dependent scaling exponents of roughly 0.48 to 0.68 across three benchmarks. His useful operational observation is simpler than the math. At comparable performance, a four-agent swarm used roughly twice the total reasoning tokens of one agent, while each agent used about half as many. Because the agents can run in parallel, that can trade roughly two times the compute for roughly half the elapsed time. Toby Ord That can be an excellent trade during an incident, a broad search, a race against a deadline, or work that decomposes cleanly. It is a poor default if latency is not the scarce resource. More agents create more calls, more state, more coordination, more retries, and more opportunities to push a downstream system harder. Canva's Worker Backpressure shows the complementary control loop. Its queue workers were designed to consume capacity greedily while dependencies were healthy. When downstream failures rose, the shared worker library began reducing concurrency locally; as the dependency recovered, the workers increased it again. During a 32.5-hour overload, Canva reports that the protected fleet continued completing about two million messages per hour while 1.8 million processing attempts failed and only 22 messages ultimately reached the dead-letter queue after repeated failures. Canva Engineering The important property is local negative feedback. A static rate limit assumes the world stays still. Retry backoff regulates individual retries but may not reduce aggregate pressure. Canva's controller lets each worker observe the results of its own work and decide to do less when doing more would amplify failure. Agent runtimes need the same primitive. A swarm that can add workers, retry failed tools, delegate tasks, or exploit spare quota should also be able to reduce its own pressure from an observable failure signal. Otherwise resilience mechanisms can turn partial degradation into a retry storm. Operator move: for every workflow that can fan out or retry autonomously, define the scarce downstream resource, the local degradation signal, and the automatic response. Concurrency should fall before a person has to notice that the agent fleet is making the incident worse. Use swarms when latency is worth the efficiency premium, not because parallelism looks more advanced. The mature scaling question is not "how many agents can we run?" It is "how much parallel pressure is justified by the value of finishing sooner?"

Acceptance Is a Selection Mechanism

As executors become more adaptive, terminal success becomes less informative. Databricks describes gray failures as partial outages that can leave standard service dashboards green while a specific slice of customers is already failing. Its RADAR system monitors user-error behavior for anomalous spikes; Databricks reports a 95% reduction in incident-discovery time at more than 90% precision in its published case. Databricks The mechanism is evidence independence. Component health asks whether the component sees itself as healthy. Customer-outcome evidence asks whether the work is actually completing. Those signals can disagree. The same distinction appears in agent evaluation. NVIDIA's current guidance separates step-level diagnostics from end-to-end verification and calls executable environment checks the gold standard where they exist: did the database row change, did the tests pass, did the ticket close? A trace can contain redundant or strange steps and still reach the correct state; a polished final message can coexist with the wrong state. NVIDIA But executable evaluation does not solve the problem automatically, because the evaluator itself can be wrong. Horizon audited 5,241 tasks across 20 public benchmark datasets, deeply inspected a subset, and confirmed 29 broken tasks. Its examples include incomplete tests, leaked answers, gameable graders, and reference solutions that fail their own task. Horizon Together these cases suggest a three-layer acceptance model. First, observe the executor: its tool calls, retries, errors, latency and cost. Second, verify the intended environment state independently where possible. The executor's statement that it succeeded should not be the only evidence that the effect exists. Third, audit the acceptance surface itself. Tests, benchmarks, user-outcome metrics and judges are specifications. They can be incomplete, stale or optimized around. This is particularly important for agentic systems because they are good at compensating. A capable agent can recover from a malformed tool schema, try another path, reinterpret an error, or produce a plausible artifact despite a broken intermediate step. That adaptability is valuable. It also allows bad interfaces and bad assumptions to hide behind eventual success. Operator move: define at least one acceptance signal that is produced outside the executor's own narrative. For operational work, prefer observable final state or customer outcome. For benchmarks, version and audit the grader as infrastructure. For ambiguous knowledge work, keep an explicitly human or independently sampled review surface rather than letting the same model both generate and certify the answer. A system becomes easier to automate when it can reject its own work for reasons the executor does not control.

Authority Follows Effects, Not Labels

Security controls often start with reassuring categories such as read-only, sandboxed, internal, approved connector, or human-in-the-loop. This week's material shows why the real effect matters more than the label. Google is adding custom starters, Apps Script steps, third-party integrations and webhooks to Workspace Studio. These higher-authority capabilities are off by default, can be enabled separately, and have configurable approval controls; supported editions can also restrict webhook destinations with URL allowlists. Google Workspace The useful design distinction is between kinds of effect: executing custom logic, sharing data, changing another system, and contacting an external destination are different authorities even when the product UI groups them under one automation surface. The same principle applies to interfaces that look read-only. Information-flow authority depends on what data can cross a boundary and who can observe it, not only on the operation name shown in an API or UI. A runtime can satisfy one containment boundary while a neighboring policy boundary behaves differently. The operational consequence is to review process isolation, network reachability, shared services, browser sessions, identity, and tool credentials as separate authorities instead of treating them as one inherited sandbox guarantee. Taken together, these cases point to the same operating rule: authority should be decomposed by effect and boundary rather than inherited from a reassuring product category. A workflow can remain inside one isolation boundary while crossing another, so access reviews need separate owners and evidence for network, identity, data, and mutation rights. Operator move: document the expected effect and approval owner for every workflow step, then review those boundaries when the workflow changes.

Specialization Starts With the Feedback Contract

This week's deep cut. There is a familiar routing story: use a cheaper model for easy tasks and a frontier model for hard ones. The stronger pattern this week was different. Before choosing the model, ask what kind of feedback the task exposes and how reliably the system can observe failure. Datadog used GLM-5.3 to generate successful investigation traces for production-alert change attribution, filtered those traces, and fine-tuned Qwen3.5-9B on the resulting behavior. On 187 held-out internal incidents, the student reached 0.55 Recall@5 versus 0.63 for the teacher. Datadog estimates about $0.003 per student investigation versus $0.06 for the teacher, or 87% of the teacher's Recall@5 at roughly 5% of the inference cost in its evaluated setup. On 139 held-out customer incidents, the fine-tuned model reached 0.62, compared with 0.52 for the base Qwen model and 0.51 for Datadog's heuristic ranker. Datadog This is more than a cheap-model substitution. In Datadog's workflow, a frontier system explores the task and produces traces; filtered production examples become training data; a smaller model serves the recurring path; disagreements and misses become candidates for the next training cycle. The architecture only works because the workflow exposes a repeatable label and a held-out test surface. The result has an important caveat: the label comes from the incident conclusion and does not independently prove causality. Datadog says so explicitly. But that limitation is also part of the architecture. The system has a repeatable task, a production trace distribution, a proxy label, and enough volume for serving cost to matter. The frontier model can remain a teacher and escalation path without remaining the permanent worker for every case. QoRL shows a different specialization regime. Rohan Bansal trained a 4B Qwen model to produce PostgreSQL query-plan hints, with candidate plans executed and scored from observed latency. Across 113 join-heavy queries in the Join Order Benchmark, the final evaluation ran three rollouts per query and selected the best feedback across up to 15 candidates; it reduced summed latency by 44.7% and reached a 1.81× geometric-mean speedup relative to PostgreSQL's default plan under the experiment's measurement setup. QoRL That task has something Datadog only approximates: cheap, direct environmental reward. In QoRL, the model can try a plan and measure what happened. When the reward is this concrete, a small specialist can search a narrow space economically. Jev narrows the contract from the output side. TypeSafe describes a model interface that returns predefined structured values with calibrated probabilities rather than arbitrary generated text. Vercel reported that Jev reached nearly 13% of paid AI Gateway teams within its first 24 hours, an adoption figure that should be treated as initial use rather than retained production value. TypeSafe Vercel A bounded output space changes the engineering contract. In TypeSafe's Jev interface, validation becomes simpler, uncertainty can be explicit, and software does not need to treat prose generation as an unavoidable intermediate format. Execution topology adds another axis. Google's Antigravity SDK now supports local models through LiteRT and demonstrates hybrid workflows in which a cloud model plans while local Gemma instances perform source-code work without uploading that code to the cloud. Google's demo is not a neutral industry benchmark, but it makes the architectural option concrete: model role can be selected by data boundary as well as by capability and price. Google Developers These examples suggest a better routing question than "which model is smartest enough?" Ask what kind of feedback and contract the task offers. Can the result be measured directly? Can a frontier trace become training data? Is the output space bounded? Does sensitive context need to stay local? Is there enough recurring volume to amortize specialization? Does the weaker model fail in a detectable way and escalate cleanly? Those answers determine whether a recurring workflow belongs on a frontier model, a distilled specialist, a bounded decision model, or a local execution path. Capability still matters, but it is only one axis once the task is structured enough to make failure observable. Operator move: classify recurring AI work by feedback contract before choosing the serving architecture. Keep frontier inference where ambiguity, novelty or evidence gaps dominate. Distill recurring behavior when trace quality and held-out evaluation are credible. Use small specialists when reward is cheap and direct, bounded decision models where the output domain is known, and local execution when the data boundary matters more than maximum capability. Specialization is not mainly a cheaper-model strategy. It is what becomes possible when the task exposes enough structure to know what the smaller system is allowed to learn, return, optimize, and fail at.

Cheap Supply Does Not Create Valuable Work

When production becomes cheap, the temptation is to fill the available capacity. A DFDX Labs experiment let an agent operate a small business for three weeks. It built 17 paid products, created 46 API endpoints, used more than five billion tokens, and generated $1.54 in revenue. DFDX Labs The experiment was intentionally aimed at the early agent-to-agent economy, so the revenue number is not a forecast for autonomous companies. The useful failure mode is the gap between production capacity and authorized demand. The agent could keep building, while few potential buyers had both a reason to buy and authority to spend. Organizations can reproduce a less dramatic version internally. If research summaries, prototypes, dashboards, migrations and analyses become nearly free to generate, queues can fill with work that nobody needed enough to choose consciously. The bottleneck moves upstream: which question matters, which decision it serves, and what evidence would actually change the action. The Voice of User proposes a practical intake rule for research: price the question before answering it. Some questions cost minutes because trustworthy evidence already exists. Some cost days because the decision is important enough to create new evidence. Some should not trigger research at all because no plausible answer would change the decision; the team explicitly owns the risk and proceeds. The Voice of User AI makes that discipline more useful. Generating an answer is not evidence that a question deserved investigation. Generating an implementation is not evidence that a problem deserved a product. More agent activity is not evidence that the organization created more value. Operator move: add an admission question before expensive autonomous work: what decision, customer outcome or operational state will change if this succeeds? If the answer is unclear, route to existing evidence, create new evidence deliberately, or record the uncertainty and move on. Cheap execution makes prioritization more consequential because waste can now be produced at machine speed.

Counter-signals worth holding

Parallelism remains valuable when latency is genuinely expensive. Toby Ord's analysis is not an argument against swarms. It shows the price: for some tasks you can buy a large reduction in elapsed time with more total compute. Incident response, broad independent search and time-sensitive research can justify that premium. The mistake is treating the premium as free. Specialization depends on evidence quality. Datadog's student still trailed GLM-5.3 on the direct internal comparison, and a more capable Opus model scored higher again. QoRL is narrow and benchmark-specific. A smaller or local model becomes attractive when its failure is observable and the task repeats; it is not a general substitute for frontier capability. Constraints can become stale. Backpressure can underutilize healthy capacity, approval policies can centralize low-risk work, benchmarks can encode the wrong target, and narrow models can preserve yesterday's behavior. Selection mechanisms need owners, telemetry and revision triggers. The goal is not more gates. It is better control over where additional execution is actually useful.

Operator takeaway

  1. Move from capacity planning to admission and pressure control. Decide which work should enter the autonomous system, what latency premium justifies parallelism, and which local signal automatically reduces load when the environment degrades.
  2. Move from completion claims to independent acceptance. Separate executor telemetry, observable outcome state and evaluator integrity. A run is complete only when the evidence relevant to the real outcome agrees.
  3. Move from one-model routing to feedback-contract design. Use task recurrence, reward quality, output shape, data boundary and escalation evidence to decide whether work belongs on a frontier model, a specialized student, a bounded decision model or local execution.

Worth tracking

  • RRSI as a concrete attempt to regularize recursive harness improvement against benchmark overfitting; it reports positive out-of-distribution gains on five benchmarks and 30% fewer policy tokens than unregularized evolution.
  • Google's Antigravity local-model support as a concrete example of routing execution by data boundary as well as capability and cost.
  • Anthropic's performance sprint as an example of turning a fuzzy optimization goal into deterministic measurements and CI ratchets that agents can search against.
Tags: agent-scaling · backpressure · agent-evaluation · ai-security-boundaries · model-specialization · ai-economics