Twitter AI Agent - 2026-08-31¶
1. What People Are Talking About¶
1.1 "Harness engineering" scaled from a taxonomy into an operational checklist, with more visible skepticism (🡕)¶
August 30 established a named prompt-context-harness-loop-graph stack; August 31 pushed the same idea into a concrete organizational checklist and drew its first clear public pushback from a credible voice. At least five items supported this theme.
@businessbarista published (180 likes, 26 replies, 19,092 views, 525 bookmarks) a 30-point list of "features of an AI native company," the highest-engagement single item in this day's enrichment set. Concrete entries included: everyone in the org using "a daily driver harness like Grok Bot, Claude Cowork, ChatGPT at Work"; a "skills distribution system" to keep agents triggering consistent skills for token efficiency; "cost per accepted PR" as a core software metric; and an "earned autonomy" ladder (observe -> suggest -> act with approval -> act alone) gated by evals. A reply from @enhansai pushed back on item 3 (a centralized queryable intelligence layer), asking pointedly where permission and meaning should actually live: "the data layer, the app layer, or the prompt?" Another reply, from @JamesSonicemi, reduced the whole list to one line: "The daily driver harness is the only line that matters. 30 features is a slide. If people still email the spreadsheet, you don't have an AI-native company."
@omarsar0 wrote (112 likes, 23 replies, 12,700 views, 44 bookmarks) that "next to evals, harness engineering is quickly becoming one of the most important skills for AI engineers to have today." A reply from @DarioGieselaar directly rejected the framing: "It's not, it's just dotfiles all over again. People go insane over it and just create problems for themselves" — one of the clearest dissents from the harness-engineering consensus seen across either day.
@RihardJarc relayed (17 likes, 3 replies, 2,921 views, 8 bookmarks) an interview with a Microsoft employee who put a number on the claim: "60% of the performance is tied to the harness and only 40% to the model." The same interview claimed Claude Code's first-pass acceptance rate beats GitHub Copilot's native pairing by an estimated 20-30%, that OpenAI models run 20-30% more expensive on complex multi-step tasks due to rework, and predicted a shift from token-based to outcome-based pricing. A reply from @unicorntrakr agreed: "Token billing was always transitional. My FD would never sign off on variable spend that depends on how verbose a model feels that day."
The exact five-layer taxonomy from August 30 (prompt -> context -> harness -> loop -> graph) recirculated again via @RoundtableSpace's post, and Andrew Ng's newly released course was widely quoted as extending the same progression one step further: prompt -> agents -> loops -> graphs -> self-improving systems, per @LunarResearcher's (68 likes, 8 replies, 6,943 views, 106 bookmarks) summary.
Discussion insight: The dissent from DarioGieselaar and the pointed governance question from enhansai mark a shift from August 30, where nearly all engagement with the harness framing was additive (naming more layers, more benchmarks). By August 31, at least some credible voices were explicitly questioning whether "harness engineering" is durable engineering discipline or repackaged tinkering.
Comparison to prior day: August 30 established the prompt/context/harness/loop/graph taxonomy and a numeric harness benchmark (Command Code's TEF bench). August 31 operationalized the same idea into an organizational checklist (businessbarista) and, notably, produced its first visible public disagreement about whether the framing is substantive.
1.2 Independent research supplied hard numbers for agent failure rates and skill persistence (🡕)¶
Where August 30's reliability discussion leaned on a DeepMind paper and an independently authored whitepaper, August 31 added two more verifiable academic papers plus a first-party update on the earlier DeepMind work, each with an image directly confirming its content.
@marfinxx summarized (18 likes, 2 replies, 1,338 views, 24 bookmarks) a Microsoft Research paper, AgentRx, which manually annotated 115 failed multi-agent trajectories and built an automated framework to localize the exact step where a run became unrecoverable. The claim that "68% of agent crashes stem from invisible cascading failures that occurred 10+ turns earlier" and specific improvement figures (root-cause localization from 31.5% to 58.7%, MTTR down 64%) are the poster's own gloss — the image confirms the paper is real (arXiv:2602.02475v1, Microsoft Research/UIUC) and does describe exactly this kind of benchmark and diagnostic pipeline, but those specific percentages are not visible in the imaged abstract and introduction.

@dair_ai reported (30 likes, 9 replies, 3,937 views, 28 bookmarks) on LoopArena, a benchmark from Alibaba's DreamX Team measuring how well a "Controller" model can guide a separate, fixed "Worker" agent through long coding tasks. The image directly confirms the headline number: "the best observed Strict Success Rate is 24.69%" on full tasks, with named failure modes — trusting a stale progress note, skipping needed verification, spending budget in the wrong direction, and stopping before the task is safe to submit.

@omarsar0 also highlighted (100 likes, 17 replies, 8,442 views, 133 bookmarks) Google's WikiSkill paper, which co-evolves agent skills with a persistent wiki-style knowledge base; the attached image confirms its central finding that evolved skills transfer across models and that "smaller models with skills can outperform substantially larger models without them." A reply from @yoav_sivan named the paper's open gap directly: "A wiki the agent can write to will slowly lie unless an entry can fail a check and get pulled. Persistence is easy. The contract is whether a stale skill is still allowed to run."

Google DeepMind's own @vivnat extended (37 likes, 4 replies, 1,941 views, 16 bookmarks) the Co-Scientist research referenced secondhand on August 30, announcing three new preprints with named academic collaborators (Genentech on cancer biology via a new PerturbME framework, Duke and Columbia on closed-loop discovery spanning materials science, biology, and computer science, and a math preprint on the sparse-literature Chowla sets problem). A reply from @PinoZlatan raised a specific, unanswered oversight question: "before each physical run, expose the exact protocol delta, affected constraint, expected information gain, and abort condition. Is that boundary evaluated, or is oversight measured only at the final decision?"
Discussion insight: Unlike August 30's reliability posts, which were largely secondhand summaries, August 31 included a first-party research-org confirmation (vivnat/DeepMind) and two independently verifiable papers with images matching their claimed content exactly (LoopArena, WikiSkill). The clearest unresolved tension across replies was not whether these methods work, but whether their proposed safeguards (constraint checks, skill validation, approval boundaries) are themselves evaluated or just asserted.
Comparison to prior day: August 30's DeepMind Co-Scientist summary came from a third party and its central "90% to 4%" hallucination-reduction figure could not be directly confirmed in the imaged page. August 31 upgraded this with a first-party DeepMind account naming specific collaborating institutions and preprints, plus two additional papers (AgentRx, LoopArena, WikiSkill) whose images directly corroborate their headline claims.
1.3 On-chain "agent commerce" volume stayed high, but split between astroturf-style promotion and more substantive analysis (🡒)¶
The termix_ai promotional cluster continued at similar volume and with the same scripted structure seen on August 30 — for example @evrendag1284's post (53 likes, 39 replies, 281 views) claiming to have personally tested the platform by sending "bids to 9 different Requests," and @trkweb3's post (19 likes, 21 replies, 96 views) restating the same identity/marketplace/reputation/verification/settlement checklist. As on August 30, several of these posts show a reply count disproportionate to their view count, consistent with coordinated reply activity rather than organic reach.
Two more substantive entries stood out against that backdrop. @NEARProtocol promoted (83 likes, 1 reply, 13,497 views, 2 bookmarks) a Delphi Digital analysis of its "agentic commerce" stack (Agent Market for bidding on open jobs, escrow-based payment, NEAR Intents for settlement), which explicitly named the unsolved problem rather than asserting it away: "deciding whether the provider delivered what was promised... becomes difficult when the work requires expertise the requester does not have." Separately, @BNBCHAIN's official account announced (62 likes, 29 replies, 19,935 views) a hackathon with "$40,000+ in prizes" to build an official BNB Agent Studio marketplace, submissions closing September 9 — a concrete, verifiable structural reason (a funded, official competition) for continued volume around on-chain agent marketplaces, distinct from the pseudonymous promotional posts.
Discussion insight: The Delphi Digital analysis is the first item across either day in this cluster to name agent-commerce's central unsolved problem (outcome verification without requester expertise) directly, rather than listing identity/reputation/settlement as if they were already solved.
Comparison to prior day: August 30 established the astroturf pattern (anomalous reply-to-view ratios, identical scripted phrasing). August 31 showed the same pattern continuing at similar volume, but also surfaced two more credible, named sources (Delphi Digital, BNBCHAIN's official hackathon) that provide non-promotional context for why the category remains active.
1.4 Coding-agent tooling shipped concrete releases addressing skill sprawl and multi-agent memory (🡕)¶
Where August 30 centered on economics and usage limits, August 31's tooling news was dominated by releases that directly answer needs raised the day before.
@VaibhavSisinty reported (19 likes, 4 replies, 3,049 views, 1 bookmark) that OpenClaw 2.0 shipped ClawHub, described as a catalog of "13,000+ community-built agent skills" installable with a single command, alongside Claude Code being able to attach to live OpenClaw sessions and OpenClaw sessions now running on remote cloud workers instead of a local machine. This is a direct, if unverified beyond the announcement, response to the skill-discovery and skill-trust problem raised on August 30 (an account there reported 2 million ungoverned GitHub agent skills with "no signal on which one was good").
@MiaAI_lab covered (47 likes, 0 replies, 2,773 views, 11 bookmarks) Hermes Agent v0.21.0, "The Pantheon Release," adding multi-agent Bot Mode with agent-to-agent DMs, persistent memory for cron jobs, and mid-task steer/stop control over subagents. A separate post in the review set put the release's scale at "5,800 commits, 2,475 merged PRs, 2,100 issues closed, 760+ contributors" since the prior version.
On the enterprise side, @mattsgarman announced (21 likes, 2 replies, 1,570 views, 5 bookmarks) that AWS Agent Registry reached general availability, naming Southwest Airlines ("went from dozens of agents with no shared record to a single governed catalog"), PepsiCo, and Syngenta as customers. A reply from @BeforeTheTurn drew a sharp distinction that applies equally to ClawHub: "A registry solves discovery; it doesn't automatically create governance. The catalog must also answer: who owns this agent, which evals keep it live, what it costs, and when it should be retired. Otherwise reuse can preserve yesterday's failure as efficiently as success."
Discussion insight: Both the consumer-facing (ClawHub) and enterprise-facing (AWS Agent Registry) responses to skill/agent sprawl arrived within a day of the problem being named, but the AWS reply thread makes clear that a searchable catalog is a necessary, not sufficient, fix — ownership, retirement, and ongoing evaluation remain open questions neither product claims to solve.
Comparison to prior day: August 30 identified skill-discovery trust and sprawl as an unmet need (2 million ungoverned skills, "which of 286 skills to actually use"). August 31 showed at least two concrete, named products (ClawHub, AWS Agent Registry) launching or reaching GA to address exactly that gap, alongside a major open-source multi-agent memory release (Hermes v0.21.0).
2. What Frustrates People¶
Loop-guided long-horizon tasks fail more than half the time, even under the best evaluated method¶
LoopArena's finding that the best observed Strict Success Rate on full, long-horizon coding tasks is 24.69% is a rare quantified confirmation of a frustration that has otherwise only been described anecdotally (context windows filling with unpruned tool output, verifiers rejecting valid work). Severity: High for anyone relying on "loop engineering" for unattended multi-round work; the named failure modes (trusting stale progress notes, skipping verification, misdirecting budget, stopping unsafely) give a concrete checklist of what to guard against. Worth building for: yes — this is a measured, not assumed, gap.
Harness-engineering discourse is starting to draw open skepticism from credible voices¶
DarioGieselaar's "it's just dotfiles all over again" reply to a well-known AI-engineering commentator, and enhansai's pointed question about where a "single source of truth" should actually live in a real system, both suggest a growing minority view that harness-engineering language is outrunning demonstrated engineering discipline. Severity: Medium — this doesn't invalidate the underlying practices, but signals that credibility of the framing itself is being tested in public for the first time.
Registries and skill catalogs solve discovery, not governance¶
The AWS Agent Registry reply (BeforeTheTurn) names a real limitation: a searchable catalog with no ownership, eval, cost, or retirement policy "can preserve yesterday's failure as efficiently as success." The same critique applies directly to ClawHub's "13,000+ skills, one command to install" framing, which addresses discovery speed but not the quality-signal problem named on August 30. Severity: Medium, likely to grow as more registries launch without governance features attached.
3. What People Wish Existed¶
Skills and knowledge that persist and transfer without silently going stale¶
WikiSkill demonstrates a working mechanism for skill persistence and cross-model transfer, but the sharpest reply (yoav_sivan) names the missing piece directly: "the contract is whether a stale skill is still allowed to run." People want not just persistent knowledge, but a validation loop that can retract a skill once it stops working. Opportunity: direct — WikiSkill is a real, published step toward this, but the validation/expiry mechanism remains unaddressed even in the paper.
Registries with real governance, not just discovery¶
Both ClawHub (13,000+ skills, one-command install) and AWS Agent Registry (a single governed catalog for Southwest, PepsiCo, Syngenta) address the same wish — findable, reusable agent components — but the AWS reply thread shows the harder ask clearly: ownership, live-eval status, cost, and retirement policy attached to every catalog entry, not just a search index. Opportunity: direct, and currently the most concretely underserved part of the registry category.
Verification of delivered work that doesn't require the requester to be an expert¶
The Delphi Digital/NEAR analysis states the agent-commerce problem precisely: escrow and identity solve payment and reputation, but "deciding whether the provider delivered what was promised... becomes difficult when the work requires expertise the requester does not have." This is a practical, urgently needed capability for any agent-to-agent marketplace to function beyond simple, easily-checked tasks. Opportunity: direct, unsolved even by the more credible actors in the space (as opposed to the termix_ai-style posts that treat it as already handled).
Diagnostic tooling that finds the real point of failure, not just the crash site¶
AgentRx's premise — that most damage happens 10+ turns before the visible crash — matches the same "verifier is the wall" complaint raised on August 30. People want automated, step-indexed fault localization rather than manual log reading after the fact. Opportunity: direct — AgentRx is a real research prototype, not yet a widely available product.
4. Tools and Methods in Use¶
| Tool | Category | Sentiment | Strengths | Limitations |
|---|---|---|---|---|
| ClawHub (OpenClaw) | Agent skill marketplace | (+) | 13,000+ community skills, one-command find/install; Claude Code can attach to live OpenClaw sessions | Discovery-focused; no evidence of quality ranking or governance beyond install mechanics |
| Hermes Agent | Open-source agent framework | (+) | v0.21.0 adds agent-to-agent DMs, persistent cron memory, mid-task subagent steer/stop; high, verifiable development velocity (2,475 merged PRs since prior version) | Open-source project; production hardening claims not independently benchmarked in this dataset |
| AWS Agent Registry | Enterprise agent catalog | (+/-) | Named enterprise adopters (Southwest, PepsiCo, Syngenta) citing reduced duplicative development | A reply notes registries solve discovery but not ownership, ongoing evaluation, cost tracking, or retirement |
| Claude Code | Coding agent | (+) | Cited as having a higher first-pass acceptance rate than GitHub Copilot (~20-30%, per one practitioner estimate); handles multi-step reasoning well per the same source | Estimate from a single paraphrased interview, not an audited benchmark |
| GitHub Copilot | Coding agent | (+/-) | Expected to close the acceptance-rate gap with Claude Code over time, per the same interview | Currently behind on first-pass acceptance per that same single source |
| LoopArena Controller pattern | Loop-engineering evaluation method | (+/-) | Isolates whether task success comes from guidance or execution by holding the Worker agent constant | Best full-task Strict Success Rate measured at only 24.69%, showing the method itself is still immature |
Sentiment toward the newest releases (ClawHub, Hermes v0.21.0) was clearly positive, framed as solving real discovery and multi-agent memory gaps. Sentiment toward the broader "harness/loop engineering" discipline itself was more mixed than on August 30, with the first visible public skepticism (dotfiles comparison) appearing this day. The clearest migration pattern is enterprises and consumer tools converging on the same shape of fix (a centralized, searchable registry) for the same underlying skill/agent sprawl problem, though neither has yet added the governance layer critics are asking for.
5. What People Are Building¶
| Project | Who built it | What it does | Problem it solves | Stack | Stage | Links |
|---|---|---|---|---|---|---|
| ClawHub | OpenClaw team | Centralized catalog of 13,000+ community-built agent skills with one-command find/install | Skill discovery and installation friction across a large, fragmented skill ecosystem | OpenClaw 2.0 platform (iOS/Android/Wear OS/web/terminal clients) | Shipped | Referenced via @openclaw announcement |
| Hermes Agent v0.21.0 | Nous Research | Open-source agent with multi-agent Bot Mode, agent-to-agent DMs, persistent cron memory, browser control | Coordinating multiple long-running agents with shared, durable memory and mid-task human control | Open-source agent framework | Shipped | hermes update |
| AWS Agent Registry | AWS | Governed enterprise catalog for publishing, discovering, and reusing internally built agents | Duplicative agent development across large organizations with no shared record | AWS enterprise platform | Shipped (GA) | Referenced via @mattsgarman |
| AgentRx | Microsoft Research / UIUC | Automated framework that pinpoints the critical failure step and category in a failed multi-agent trajectory | Manual log inspection is too slow and unreliable to localize root causes in long, multi-agent runs | Constraint-generation + step-indexed validation pipeline | RFC (research preprint) | arXiv:2602.02475v1 |
| LoopArena | Alibaba DreamX Team / UNSW Sydney / CSIRO | Benchmark evaluating how well a Controller model can guide a separate Worker agent through long coding tasks | No existing way to separate loop-guidance quality from the underlying coding agent's own ability | Controller/Worker two-agent harness, benchmark + eval code | RFC (research preprint, code released) | github.com/AMAP-ML/LoopArena |
| WikiSkill | Google Research / Virginia Tech | Co-evolves agent skills alongside a persistent wiki-style knowledge base for cross-model transfer | Skill-tuning insights getting scattered and lost across optimization iterations | Persistent knowledge base + skill-evolution framework | RFC (research preprint) | arXiv:2608.27454v1 |
ClawHub and AWS Agent Registry both illustrate the same build pattern from two different markets (consumer/developer vs. enterprise): centralize a fragmented, duplicative skill or agent inventory behind one searchable, installable catalog. The three research artifacts (AgentRx, LoopArena, WikiSkill) all target different parts of the same underlying reliability gap — diagnosing failures after the fact, measuring loop-guidance quality, and persisting/evolving skills — suggesting agent reliability tooling is being attacked from multiple independent research groups simultaneously rather than converging on one approach yet.
6. New and Notable¶
AgentRx gives multi-agent failure diagnosis a measurable, automated method¶
Microsoft Research's AgentRx (arXiv:2602.02475v1) replaces manual log inspection with a constraint-generation and step-indexed validation pipeline to localize the exact turn where a multi-agent run became unrecoverable, evaluated across 115 manually annotated failed trajectories.
LoopArena puts a hard number on "loop engineering" maturity¶
Alibaba's LoopArena benchmark (arXiv:2608.28281v1) measured the best evaluated Controller model at a 24.69% Strict Success Rate on full long-horizon tasks — a concrete ceiling on how well current models can reliably guide another agent through a long coding task today.
WikiSkill shows persistent knowledge, not bigger models, drives skill transfer¶
Google's WikiSkill paper (arXiv:2608.27454v1) demonstrates that skills evolved with a persistent wiki-style knowledge base transfer across model families, and that smaller models with evolved skills can outperform larger models without them.
ClawHub and AWS Agent Registry both launched as direct responses to skill/agent sprawl¶
Within a day of the skill-discovery and trust problem being named in this dataset (2 million ungoverned GitHub skills), both a consumer-facing skill marketplace (ClawHub, 13,000+ skills) and an enterprise agent catalog (AWS Agent Registry, GA, with Southwest Airlines/PepsiCo/Syngenta as named customers) shipped centralized discovery mechanisms.
7. Where the Opportunities Are¶
[+++] Automated failure diagnosis and loop-guidance evaluation for multi-agent systems — AgentRx and LoopArena both supply hard, published numbers (68% of crashes from early cascading failures per AgentRx's claimed results; only 24.69% Strict Success Rate on full loop-guided tasks per LoopArena) showing this is a measured, unsolved gap rather than a vague complaint, with two independent research groups (Microsoft, Alibaba) attacking it simultaneously.
[++] Governance layer on top of agent/skill registries — ClawHub and AWS Agent Registry both shipped discovery mechanisms this day, but the AWS reply thread (BeforeTheTurn) names the still-open need precisely: ownership, live-eval status, cost tracking, and retirement policy per catalog entry, not just search.
[++] Persistent, validated skill knowledge bases — WikiSkill shows skill evolution and cross-model transfer work, but the sharpest reply (yoav_sivan) identifies the missing validation/expiry contract as the actual product opportunity, not the persistence mechanism itself.
[+] Outcome verification for agent-to-agent commerce — The NEAR/Delphi Digital analysis is the first source across both days to name this gap explicitly rather than assume it away; the surrounding termix_ai-style promotional volume suggests strong perceived demand for this category even though a working, expertise-independent verification mechanism has not yet been demonstrated.
8. Takeaways¶
- "Harness engineering" discourse drew its first visible, credible public pushback ("it's just dotfiles all over again"), a shift from August 30's near-uniformly additive coverage of the same framing. (DarioGieselaar reply to omarsar0)
- Two independent, verifiable research papers quantified the reliability gap directly: AgentRx traces most multi-agent crashes to failures many turns before the visible break, and LoopArena measured the best current loop-guidance method at a 24.69% full-task success ceiling. (dair_ai)
- Google DeepMind's own account extended and named collaborators for the Co-Scientist research referenced secondhand on August 30, adding math, cancer-biology, and closed-loop discovery preprints with Genentech, Duke, and Columbia. (vivnat)
- Two concrete products (ClawHub, AWS Agent Registry) launched within a day of the skill/agent sprawl problem being named, but a reply thread on the AWS launch made clear that search-based discovery is not the same as governance (ownership, eval status, retirement). (mattsgarman)
- On-chain agent-commerce volume remained high and showed the same coordination markers as August 30, but the most substantive entry (NEAR/Delphi Digital) was notable for naming its central unsolved problem — outcome verification without requester expertise — directly, rather than asserting it solved. (NEARProtocol)