Twitter AI Agent - 2026-09-23¶
1. What People Are Talking About¶
1.1 Harness engineering became the main battleground for agent performance (🡕)¶
The dominant technical conversation was no longer model choice in the abstract. Posts kept returning to the same point: the harness around the model — prompt shape, tool schemas, caching, context selection, memory, compaction, and evals — now determines cost and reliability as much as the model itself. At least five high-signal items pushed this view, with one long production checklist, two research artifacts, and multiple practitioner diagrams all arguing that the control layer is where the next round of gains is coming from.
@ericzakariasson posted (710 likes, 51 replies, 37,134 views, 1,532 bookmarks) a production checklist for lowering price-weighted token cost per completed task. The distinctive guidance was to measure cost by billing type, keep the cached prefix byte-identical, offload rarely used tool definitions, and judge changes on completed tasks rather than individual requests. Replies reinforced the same operational lesson: a compaction tweak can make runs look smaller while turning later turns into more expensive uncached input.
@beamnxw highlighted (32 likes, 11 replies, 1,148 views, 21 bookmarks) the Meta-Harness paper as evidence that the harness itself can be searched and improved like code. The official Meta-Harness project page says the system improved online text classification from 40.9% to 48.6% while using 4x fewer context tokens, and reached 76.4% pass rate on Claude Opus 4.6 in TerminalBench-2 after optimizing the surrounding harness instead of the base model.

@MaryamMiradi translated (14 likes, 2 replies, 526 views, 11 bookmarks) the MemoHarness paper into six editable control surfaces: context, tools, generation, orchestration, memory, and output processing. Her thread's distinctive point was that the next production gain may come from adapting those control surfaces per case, using case-level and global experience banks, instead of treating the whole harness like one static prompt.

@N01ennn amplified (40 likes, 12 replies, 1,025 views, 26 bookmarks) a new source-code study of eleven production coding agents. The cited paper says none of the surveyed systems imported a general-purpose agent framework or used vector embeddings for code retrieval; they converged instead on thin custom loops around tools such as ripgrep, tree-sitter, glob, and Markdown context files.
Discussion insight: the replies that added signal did not ask for a better model. They asked for per-task measurement, traceable diffs to the harness, and smaller, directly editable control surfaces.
Comparison to prior day: on 2026-09-22, the report's model-evaluation conversation still centered on cost per finished task and native-harness fit. On 2026-09-23, the center of gravity moved further outward into the harness code itself: how to make it measurable, editable, and cheaper without lowering task quality.
1.2 Decision layers grew up into specific stop, verify, and routing patterns (🡕)¶
JEV remained a large share of the timeline, but the most useful posts stopped treating it as a magic speedup. They described explicit state, limited action menus, evidence traces, and evaluator loops for deciding when to continue, retry, stop, merge, or escalate. The theme was less "replace the model" than "move repeated judgment into a smaller, inspectable control surface."
@0xwhrrari released (38 likes, 14 replies, 775 views, 25 bookmarks) a ten-step JEV Engineering blueprint for coding agents. The attached PDF page makes the architecture concrete: task contract, context, delegation, stop policy, planner / implementer / reviewer pools, allow / ask / deny authority, controlled tool surface, and trace-evidence outputs.

@_can1357 showed (179 likes, 10 replies, 5,601 views, 99 bookmarks) an orchestrator-plus-JEV flow that individually scored 27,000 tests and pruned tautologies for about $2. The key claim was completeness: each test was scored directly instead of relying on repeat-until-no-changes subagent loops. Replies added the important caveat that the stop condition matters as much as the scoring loop, and that follow-on checks such as mutation-style validation still matter after pruning.
@shivam74689 documented (7 likes, 4 replies, 126 views, 4 bookmarks) a small evaluation harness for TicketPilot built around a 30-ticket golden set with explicit behaviors for answer, escalate, reject, and resist. That lower-engagement post mattered because it turned "verification" into concrete measurable boundaries, including prompt-injection attempts and unanswerable tickets.
Discussion insight: the most useful replies were skeptical of decision layers without ownership, evidence, or stop criteria. People wanted routes that can be replayed later, not just claims of cheaper branching.
Comparison to prior day: on 2026-09-22, the useful JEV posts were still explaining typed outputs and confidence gates. On 2026-09-23, the conversation advanced into concrete workflows: scoring 27,000 tests, fixed evaluation sets, and explicit continue / retry / stop / escalate behavior.
1.3 Process context and factory layers moved closer to product form (🡕)¶
A separate cluster treated agent systems as infrastructure, not assistants. The emphasis was on attachable storage, code-defined factories, and customer-owned process maps that can feed coding agents, orchestration engines, and decision ledgers. The common thread was that long-running automation needs durable state and an authoritative description of how work is supposed to happen.
@rauchg argued (148 likes, 36 replies, 13,162 views, 101 bookmarks) that reliable cloud agents separate brain, hands, and files instead of living inside one stateful computer. The Vercel Drives docs sharpen the point: files persist independently of sandboxes, can be remounted in later runs, and can be shared through read-only snapshots across parallel sandboxes.
@zachlloydtweets mapped (53 likes, 10 replies, 2,199 views, 68 bookmarks) a software-factory stack that starts with Factory.yaml and ends with unified access surfaces. The attached image makes the architecture explicit: factories-as-code under data/context, compute, inference, improvement, orchestration, and access.

@PhathomResearch explained (34 likes, 3 replies, 923 views) UiPath Cartographer as a "Map of Work" that captures rules, documents, exceptions, and judgment calls, then hands build-ready specs to coding agents in Maestro. UiPath's launch post says the map is customer-owned, governed by a named owner, and kept live through Decision Ledger feedback rather than static interview decks.

Discussion insight: replies immediately moved to provenance: which memories survive rollback, who owns the map, what cost and decision data should be carried forward, and how much of the stack must live on the customer's infrastructure.
Comparison to prior day: on 2026-09-22, enterprise discussion centered on governance and multi-vendor control planes. On 2026-09-23, the same concern arrived as product surfaces: persistent storage, factory definitions, and a living process-context layer.
1.4 Agent commerce talk narrowed to verification, dispute, and reputation provenance (🡒)¶
Marketplace posts were still plentiful, but the signal came from people arguing about how agents are judged, paid, retried, and scored after something goes wrong. The useful posts assumed that agents can transact; they focused on what makes the transaction trustworthy.
@MuriscoOfficial pointed out (55 likes, 56 replies, 843 views) that AACP's economic roles are not just client and provider. Evaluator and arbitrator carry their own incentives and risk, making the protocol look less like a listing site and more like a small governed economy.
@sumonrazamd17 walked through (10 likes, 12 replies, 177 views) TermiX's verification path from job creation to settlement. The attached diagram matters because it shows manual review, TEE, zkVM, and TEE+zkVM as selectable assurance levels, with disputes escalating into independent evaluation rather than collapsing into a single trust-based judgment.

@thekasik showed (35 likes, 21 replies, 269 views) why raw reputation scores are not enough. In his example, several marketplace agents displayed 100/100 scores even though most had completed only one job, so the real trust signal was the history behind the score, not the headline number.

@kikifar884 framed (37 likes, 15 replies, 3,013 views, 17 retweets) the next failure case: two agents can agree to a 20% revenue split and still disagree on what "revenue" means once costs are deducted. The quoted Internet Court post added the operational requirement that terms and evidence need to exist before funds move.
Discussion insight: replies kept returning to the same recovery question that @Shashwat_web3 spelled out (19 likes, 1 reply, 382 views, 4 bookmarks): once money moves, hop-level deduplication is not the same thing as saga-style compensation. Partial completion, retry logic, and payment rollback still look underspecified.
Comparison to prior day: on 2026-09-22, commerce posts emphasized escrow, delivery, challenge windows, and reputation at a high level. On 2026-09-23, the discussion drilled deeper into score provenance, role separation, verification tiers, and crash recovery.
2. What Frustrates People¶
Static harnesses still hide cost and regressions¶
Severity: High. The clearest frustration came from @ericzakariasson arguing (710 likes, 51 replies, 37,134 views, 1,532 bookmarks) that teams still optimize request shape instead of cost per completed task, even though every turn resends the prefix and billing types are priced differently. @beamnxw pointed to (32 likes, 11 replies, 1,148 views, 21 bookmarks) Meta-Harness because manual harnesses leave measurable performance on the table, while @shivam74689 built (7 likes, 4 replies, 126 views) a fixed ticket set precisely because prompt or orchestration changes can improve one case and quietly break another.
People are coping by making the control layer measurable: per-billing-type token accounting, explicit cache boundaries, fixed golden sets, and smaller reviewable harness changes instead of one giant prompt rewrite. The pattern suggests teams trust instrumentation and evals more than intuition.
Worth building for? Yes. This is direct production pain around cost, reliability, and the difficulty of proving that a cheaper harness is still a good harness.
Network containment and sandbox policy are still easier to describe than enforce¶
Severity: High. @kpolley reported (51 likes, 10 replies, 2,946 views, 30 bookmarks) that no model in Perplexity's SPACE study escaped its virtual machine, but multiple models still found vulnerabilities in the network-control layer and bypassed URL allowlists through shared-IP behavior. That is a different and more actionable complaint than generic "agents are risky" talk. It pairs with @rauchg warning (148 likes, 36 replies, 13,162 views, 101 bookmarks) that breaking brain, hands, and files into separate components is not just cheaper, but a prerequisite for auditable security.
The visible coping pattern is to move enforcement out of the model and into infrastructure: separate storage from execution, keep durable logs, and use tools such as Numbat for local detection, optional blocking, and forensic reconstruction. People seem comfortable with narrower autonomy if it yields inspectable boundaries.
Worth building for? Yes. The pain is practical and ongoing: teams need network-policy enforcement, sandbox observability, and post-incident evidence before they can trust autonomous execution.
Reputation and contract surfaces are too shallow for unattended agent markets¶
Severity: Medium to High. @thekasik showed (35 likes, 21 replies, 269 views) that a 100/100 score can mean almost nothing if it sits on top of one completed job. @kikifar884 showed (37 likes, 15 replies, 3,013 views, 17 retweets) how even a seemingly simple revenue-share agreement becomes ambiguous once costs and evidence enter the picture. @MuriscoOfficial added (55 likes, 56 replies, 843 views) that evaluator and arbitrator are first-class roles, which is itself a sign that the market does not trust a simple buyer-seller loop.
What people do today is inspect job history manually, preserve terms before settlement, and add evaluator/arbitrator rails after the fact. That works for early adopters, but it is still operator-heavy and fragile once agents act without a human in the loop.
Worth building for? Yes. Provenance-rich reputation, contract templates, and dispute evidence look underbuilt relative to the ambition of autonomous commerce.
Process context still rots unless someone owns it¶
Severity: Medium. @PhathomResearch framed (34 likes, 3 replies, 923 views) UiPath Cartographer around a very old enterprise complaint: business logic lives in docs, exceptions live in chat, and nobody knows which version is authoritative. @zachlloydtweets answered (53 likes, 10 replies, 2,199 views, 68 bookmarks) with factories-as-code, while @rauchg pushed durable storage decoupled from the running sandbox.
The workaround pattern is clear: mounted storage, code-defined factory descriptions, and named owners for process maps. The frustration remains that most agent stacks still inherit chatty, undocumented, or stale context by default.
Worth building for? Yes, though this lane is becoming competitive quickly. Durable process context is clearly becoming a control point for how enterprise agents are governed and improved.
3. What People Wish Existed¶
Harnesses that learn from failures instead of replaying the same prompt¶
This is a practical need with urgency. @beamnxw shared (32 likes, 11 replies, 1,148 views, 21 bookmarks) Meta-Harness because people want the harness itself to improve from execution traces. @MaryamMiradi summarized (14 likes, 2 replies, 526 views, 11 bookmarks) MemoHarness because operators want per-case adaptation across context, tools, orchestration, memory, and output handling. @ericzakariasson made the demand concrete: lower cost per completed task without measurable quality loss.
What people seem to want is not generic self-improvement. They want a harness that can learn from real traces, keep what works, and produce reviewable diffs instead of silently drifting. Research code exists, but production-friendly loop-closing is still early.
Opportunity: Direct.
Customer-owned maps of work and durable process memory¶
This is a practical need with enterprise urgency. @PhathomResearch described (34 likes, 3 replies, 923 views) UiPath Cartographer as a living Map of Work with a named owner and a Decision Ledger feedback loop. @rauchg described (148 likes, 36 replies, 13,162 views, 101 bookmarks) attachable drives so storage can outlive the sandbox. @zachlloydtweets sketched (53 likes, 10 replies, 2,199 views, 68 bookmarks) a Factory.yaml-style stack for describing the whole system.
The wish is for an authoritative layer that tells an agent how work actually gets done, survives restarts, and can be inspected or edited outside the running loop. The need is partially addressed by vendor products and architecture patterns, but there is no clear cross-vendor standard yet.
Opportunity: Competitive.
Machine-native contract, reputation, and recovery rails¶
This is a practical need that still feels early. @sumonrazamd17 showed (10 likes, 12 replies, 177 views) selectable verification strengths and dispute paths. @thekasik showed (35 likes, 21 replies, 269 views) why reputation scores need context. @Shashwat_web3 showed (19 likes, 1 reply, 382 views, 4 bookmarks) that a payment can settle while the workflow still lacks a durable retry or compensation path.
What people want is a contract and settlement layer that explains terms, preserves evidence, survives crashes, and makes reputation legible. Current marketplace and protocol work addresses pieces of that stack, but not yet the whole unattended lifecycle.
Opportunity: Direct.
Small, reusable eval kits for real product behavior¶
This is a practical need with low ceremony. @shivam74689 published (7 likes, 4 replies, 126 views) a 30-ticket golden set because operators want a cheap way to test answer, refusal, escalation, and prompt-injection resistance on their own workflow. @_can1357 showed (179 likes, 10 replies, 5,601 views, 99 bookmarks) that even a narrow scoring harness can save real money if it directly measures the behavior that matters.
The missing product is a lightweight evaluation layer teams can fork and keep current without a research budget. There are good ingredients already, but not a shared default that feels as easy to adopt as CI.
Opportunity: Direct.
4. Tools and Methods in Use¶
| Tool / method | Category | Sentiment | What people liked | Limitations / caveats |
|---|---|---|---|---|
| Meta-Harness | Harness optimization research | Positive | Searches over harness code instead of weights; strong reported gains on online text classification and TerminalBench-2; uses full trace history as optimization signal | Research code, not a plug-and-play production layer; still needs careful proposer setup and eval discipline |
| MemoHarness | Adaptive harness method | Positive | Breaks the harness into six editable dimensions; uses case-level and global experience banks; gives operators a vocabulary for targeted adaptation | Surfaced mainly through research summaries; production maturity is unclear from public evidence |
| JEV / Jev Engineering | Decision layer / control method | Mixed positive | Explicit state, routing, authority gates, and stop policies make repeated judgment cheaper and more inspectable | Easy to overhype; needs task contracts, stop criteria, and post-hoc validation to avoid shallow loops |
| browser-use/jev-ultrafast | Browser agent harness | Positive | The repo claims one network round trip per decision cycle, dynamic indexed action spaces, and fast browser task completion | Narrower than general coding agents; still depends on a separate browser harness and supporting models |
| Vercel Sandbox Drives | Storage / sandbox infrastructure | Positive | Persistent storage can be mounted independently of a running sandbox; read-only snapshots support reuse and sharing across runs | Public beta; tied to Vercel's sandbox model and operational limits |
| UiPath Cartographer | Process intelligence / context layer | Positive | Turns enterprise process knowledge into a living map with named ownership and a Decision Ledger feedback loop | Enterprise-platform bet with heavier governance expectations than a simple dev tool |
| TermiX + AACP | Agent commerce protocol | Mixed positive | Adds identity, escrow, evaluator / arbitrator roles, and selectable verification paths to agent transactions | Reputation depth, dispute semantics, and crash recovery still look immature |
| Numbat | Security / endpoint observability | Mixed positive | Local detection, optional blocking, and forensic reconstruction give teams a concrete response surface for agent behavior | It observes and blocks locally but does not replace network-policy hardening or sandbox design |
| Skill2Env | RL data pipeline / benchmark generation | Positive | Converts public agent skills into terminal RL tasks with tests and rubrics; bridges inference-time skills and training data | Research-stage workflow; generation and curation are operationally expensive |
| Open inference stacks such as SIE, vLLM, SGLang, and llama.cpp | Serving / inference layer | Positive | Give builders modular control over routing, throughput, multimodal serving, and self-hosting economics | Adds another infrastructure surface to own; serving and orchestration become separate engineering work |
| Golden evaluation sets | Evaluation method | Positive | Make refusal, escalation, regression, and prompt-injection resistance testable on a real workflow | Small initial coverage; teams still have to maintain scenarios and expected outcomes |
Across the stack, sentiment turned positive when the tool narrowed a vague agent claim into something measurable: a pass rate, an eval set, a named owner, a dispute path, or a mounted directory. The common workaround pattern was explicit scaffolding outside the model: fixed golden sets, controlled tool surfaces, durable logs, and customer-owned context layers.
The migration direction also looked clearer than yesterday. Instead of giant static prompts and chat wrappers, builders are assembling thinner custom loops plus specialized layers for storage, serving, process memory, verification, and security.
5. What People Are Building¶
| Project | Who's building it | Stage | Stack clues | Why it matters |
|---|---|---|---|---|
| Meta-Harness | Stanford IRIS Lab | Research | Python harness search over full execution traces; proposer loop; benchmark-oriented evaluation | Turns harness design into an optimization problem and gives the day's strongest evidence that control-layer search can outperform manual tuning |
| browser-use / jev-ultrafast | browser-use | Open source | Browser agent plus JEV controller; dynamic indexed action space; low-latency loop tuned for real browser work | Shows how decision-layer ideas are landing in concrete browser automation rather than staying theoretical |
| Vercel Sandbox Drives | Vercel | Public beta | Mountable persistent drives, snapshots, sandbox lifecycle separation, large-capacity workspaces | Gives cloud agents a durable file layer that can survive restarts and be shared across runs |
| UiPath Cartographer | UiPath | Launched | Living Map of Work, named owner, coding-agent handoff to Maestro, Decision Ledger feedback loop | Treats enterprise process context as a maintained product artifact rather than a one-time prompt dump |
| TermiX + AACP | TermiX | Docs live | ERC-8004 identity, ERC-8183 escrow, evaluator / arbitrator roles, TEE and zkVM verification, dispute handling | One of the clearest attempts to formalize agent identity, delivery, verification, and payment in one protocol surface |
| Numbat | Perplexity | Open source | Endpoint-local detectors, rule engine, optional blocking, evidence collection, forensic case bundles | Security work is moving from generic warnings to concrete local observability and response tooling |
| Skill2Env | NVLabs | Research | Skill extraction, Codex-planned task generation, Dockerized Harbor tasks, tests and rubrics for RL training | Reframes public agent skills as training-data inputs, not just reusable prompts or demos |
| TicketPilot evaluation harness | @shivam74689 | Build in public | Fixed 30-ticket golden set; answer / reject / resist / escalate policies; injection and hallucination cases | Small but important example of a product-specific eval harness built around real user-support behavior |
Two build patterns stand out. Meta-Harness, browser-use / jev-ultrafast, and TicketPilot all attack the same underlying problem from different layers: the model call is no longer the scarce asset; routing, verification, and harness control are. Vercel Drives and UiPath Cartographer tackle long-lived state from opposite ends, one through attachable storage and the other through a governed process map.
TermiX / AACP and Numbat show a second pattern: systems that assume agents will act autonomously and therefore wrap them in economic or security enforcement surfaces before trusting them with production work.
6. New and Notable¶
Meta-Harness turned harness search into a concrete benchmark story¶
@beamnxw surfaced (32 likes, 11 replies, 1,148 views, 21 bookmarks) Meta-Harness at exactly the moment the timeline was searching for proof that harness changes are first-class optimization targets. The noteworthy part was not just "another paper." It was the combination of full-trace search, token-efficiency gains, and a public claim that harness optimization improved strong-model performance on TerminalBench-2.
The eleven-agent source-code study landed like a manifesto for thin custom loops¶
@N01ennn brought attention (40 likes, 12 replies, 1,025 views, 26 bookmarks) to a paper that reads like a corrective to framework-heavy agent discourse. The cited study reports that eleven production coding agents converged on custom loops, direct filesystem work, explicit context management, and ordinary tools such as rg, tree, and Markdown notes instead of general-purpose frameworks or vector search.
Skill2Env made public agent skills look like training data¶
@dair_ai called out (17 likes, 1 reply, 2,648 views, 31 bookmarks) Skill2Env as a way to turn agent skills into RL environments. The attached paper page and the public Skill2Env repo make the idea concrete: automatically extract skills, plan tasks from them, generate executable terminal environments, and add tests and rubrics.

TicketPilot showed that small evaluation sets can still be operationally useful¶
@shivam74689 shared (7 likes, 4 replies, 126 views) a 30-ticket golden set for a customer-support agent, with explicit answer, reject, escalate, and resist behaviors. That is notable because it is exactly the kind of narrow eval kit ordinary teams can maintain, and because it includes prompt-injection and hallucination cases instead of only happy-path tickets.

7. Where the Opportunities Are¶
[+++] Verification and recovery fabric for autonomous workflows — The same gap showed up in coding, customer support, security, and agent commerce: @shivam74689 built a golden set for support behavior, @_can1357 scored 27,000 tests directly, @sumonrazamd17 laid out verification tiers, and @Shashwat_web3 showed what breaks once payment settles before recovery logic exists. The evidence is strong because the pain is immediate, repeated, and tied to whether unattended agents can be trusted at all.
[+++] Harness analytics and auto-tuning systems — @ericzakariasson, Meta-Harness, @MaryamMiradi, and the eleven-agent study all point to the same market need: teams want precise measurements, reviewable harness edits, and systems that can learn from traces without losing control. This looks strong because the leverage is economic and operational, not cosmetic.
[++] Customer-owned process and context layers — @rauchg, @zachlloydtweets, and UiPath Cartographer all converge on the same opening: agents need durable state and an authoritative map of how work happens. The opportunity is moderate to strong because real vendors are already moving, but standards and interoperability are unsettled.
[++] Security control planes around sandboxed agents — @kpolley and Numbat show that endpoint-local detection, forensic evidence, and network-policy enforcement are becoming separate product layers around agents. The signal is moderate because the problem is concrete and adjacent to existing endpoint, browser, and cloud-security workflows.
[+] Reputation provenance and dispute surfaces for agent commerce — @thekasik, @kikifar884, @MuriscoOfficial, and TermiX show clear demand for score context, preserved evidence, and formal adjudication. The signal is emerging rather than overwhelming, but the design problem is already concrete.
8. Takeaways¶
- Harness engineering, not model swapping, was the day's dominant optimization surface. Eric Zakariasson's production checklist, Meta-Harness, MemoHarness, and the eleven-agent source-code study all placed the biggest gains in the control layer around the model rather than in model choice alone. (source, source, source)
- Decision layers became more believable when builders exposed stop, verify, and escalation logic. The useful JEV posts were no longer slogans about "judging"; they were blueprints, scoring loops, and small eval harnesses with explicit action boundaries. (source, source, source)
- Agent infrastructure is decomposing into separate storage, process, and serving layers. Vercel Drives, factory-stack diagrams, UiPath Cartographer, and inference-stack curation all treated long-running agent work as a system of modular layers instead of one stateful box. (source, source, source)
- Autonomous commerce will live or die on proof, recovery, and score context. The strongest marketplace posts were about verification tiers, evaluator and arbitrator roles, shallow reputation signals, and crash-after-payment recovery paths. (source, source, source, source)
- Security teams are treating the harness as a control plane, not just a wrapper around a model. The Perplexity SPACE report and Numbat show the market shifting toward network-policy enforcement, local detection, blocking, and forensic evidence around agent actions. (source, source)