Skip to content

Twitter AI Agent - 2026-08-05

1. What People Are Talking About

1.1 Skills became installable operating layers, not just reusable prompts (🡕)

The strongest packaging signal was a shift from skill design to skill distribution. Four retained items focused on docs, installers, official plugin channels, curated bundles, and operating layers that sit above individual skills. Compared with August 4's discussion of skills as evidence-bearing artifacts, August 5 pushed the conversation toward how people discover, install, route, and maintain them in day-to-day agent work.

@mattpocockuk released (1,219 likes, 49 replies, 38,104 views, 747 bookmarks) mattpocock/skills v1.2 with full docs, a Claude Code marketplace plugin, Codex support via agents/openai.yaml, and new workflow skills such as /wizard, /to-questionnaire, and /wait-what. The attached installer page is informative because it shows the two public distribution paths side by side: editable npx skills@latest add mattpocock/skills installs and a managed Claude Code plugin, while also surfacing the project's 13.5 million installs and cross-agent positioning.

AIHero installer page showing editable installs, Claude Code marketplace install, and 13.5 million skills.sh installs

The public README reinforces that packaging strategy: one path copies skills into a repo for local editing, while the plugin path keeps them managed and read-only. A reply added the day's most direct pain signal: the skill set is useful precisely because it is broad, but broad enough to feel overwhelming without a helper that recommends the right skill for a job.

@rlaope argued (45 likes, 1 reply, 3,317 views, 60 bookmarks) that Hermes users “really only need ONE plugin,” then positioned oh-my-hermes as an operating layer above Hermes-native skills. Its public site says it adds planning, research, coding handoff, operations, and project memory without replacing Hermes; the attached images matter because they show the same workflows spanning Hermes Desktop, CLI, and messenger, plus a live issue backlog on memory tiers, scoped receipts, reviewed skill drafts, and schedule policy rather than a static launch poster.

@alex_prompter proposed (8 likes, 2 replies, 3,443 views) a compact SKILL.md anatomy with explicit triggers, inputs/outputs, and “NOT FOR” boundaries. That was a smaller signal, but it fit the same trend: people are trying to make skills auditable and composable enough that large libraries do not collapse into routing ambiguity.

Discussion insight: The useful disagreement was not whether skills work, but how people keep them usable. Matt Pocock's replies surfaced discovery overload, while oh-my-hermes explicitly sold itself as a cure for plugin fatigue and misrouted workflows.

Comparison to prior day: August 4 framed skills as trainable and evidence-backed. August 5 shifted toward packaging, install surfaces, and the operating layers required to keep growing skill catalogs usable.

1.2 Harness engineering became a public benchmark and reverse-engineering race (🡕)

Five retained items treated harnesses as the real product surface: something to document, reverse engineer, benchmark, and co-train against. The theme expanded August 3's measurement vocabulary and August 4's lifecycle conversation into a more public competition over who can best expose, explain, and improve the system around the model.

@morganlinton bookmarked (505 likes, 3 replies, 129,314 views, 787 bookmarks) Pi's harness-engineering guide as “the single best thing” he had read on efficient harness engineering, which was itself evidence that the topic had become worth reading as a discipline. More concretely, @swyx pointed (82 likes, 10 replies, 8,623 views, 83 bookmarks) readers to Latent.Space's public reconstruction of ChatGPT Work, which describes a Codex-harness-based knowledge-work product running in cloud microVMs with browser use, plugins, and product-managed cross-task services such as Library, Projects, Personal Context, and Memory. The attached context-architecture image is informative because it makes one non-obvious design choice explicit: task-local working directories sit below a separate OpenAI-managed layer, so cross-thread continuity is not just “shared files on a computer.”

ChatGPT Work context architecture showing task-local working directories beneath an OpenAI-managed layer for library, projects, plugins, personal context, and memory

@johannes_hage introduced (152 likes, 5 replies, 8,107 views, 51 bookmarks) Prime Agent, whose public launch post describes a persistent IPython REPL, programmatic sub-agent calling, direct agent-to-agent messaging, and CRUD over prompts, memories, skills, and subagents. The architecture diagram matters because it shows the claimed control loop directly: task -> model -> IPython kernel -> rlm / sub-agents -> continual harness state.

Prime Agent architecture diagram showing model, persistent IPython kernel, sub-agents, and continual harness state

A second Prime Agent tweet from @Dr_Singularity amplified (343 likes, 21 replies, 20,024 views, 86 bookmarks) the benchmark side of the story. Its chart is informative because it shows Prime Agent + Opus 5 reaching 95.5% on ARC-AGI-3, slightly above the 95.4% human baseline line, with separate curves for Sol, Terra, and GLM 5.2.

ARC-AGI-3 compute-scaling chart showing Prime Agent plus Opus 5 reaching 95.5 percent versus the 95.4 percent human baseline

@ren_hongyu released (166 likes, 9 replies, 8,633 views, 8 bookmarks) Muse Spark 1.2 and Muse Code, which Meta says is a terminal coding agent with persistent async background agents, replay-exact event logs, and bundled /plan, /grill, and /goal skills. The benchmark image compared Muse Code against Claude Code, Codex, Grok Build, and others across Terminal-Bench 2.1, DeepSWE 1.1, Meta Internal Coding Bench, GDPVal-AA V2, and MCP Atlas, but the replies added useful skepticism: one reader said Muse Code was hard to discover on the public web, and another challenged the honesty of the charting.

Discussion insight: The biggest credibility pattern was that third-party writeups are becoming part of the product surface. Replies under the Latent.Space post said harness breakdowns are often more useful than official docs, while Muse Code's replies showed that launch charts without transparent public surfaces trigger immediate pushback.

Comparison to prior day: August 3 treated the harness as something to benchmark. August 4 connected it to traces, memory, and lifecycle controls. August 5 turned it into a public teardown and launch category in its own right.

1.3 Owning the post-build surface mattered as much as building the agent (🡕)

A third cluster focused on everything that happens after the core agent exists: who can use it, how it gets discovered, what permissions it carries, and how it pays its way. Four retained items made the same point from different angles: shipping the model loop is not enough if access, routing, distribution, and settlement remain improvised.

@rileybrown said (71 likes, 8 replies, 7,636 views, 100 bookmarks) Vercel's internal agent V already works alongside almost 1,000 employees, then structured the public conversation around a familiar fork: one GOD agent or a team of smaller agents with different permissions and computer access. The most useful replies cut both ways. One argued that permission boundaries eventually force the multi-agent answer; another warned that a single internal agent can look stronger than it really is because it only sees “Vercel-shaped work.”

@Fetch_ai summarized (107 likes, 1 reply, 6,067 views) the post-build surface as deployment, discovery, orchestration, user access, and payments, then mapped its stack across uAgents, Agentverse, ASI:One, and Fetch Business. In parallel, @ama_protocol positioned (54 likes, 49 replies, 611 views) Amadeus as both a destination product (ama hub) and an embedded agent layer (AMA Embed) for third-party apps, with private-but-provable records around real-money actions.

@wardenprotocol launched (47 likes, 14 replies, 7,177 views) Halo, a public-alpha peer-to-peer AI inference marketplace on Base where agents can act as first-class consumers, pay from USDC deposits, and use local OpenAI-compatible endpoints without carrying provider API keys. The launch post's clearest claim is not model novelty but distribution and settlement design: wallet as credential, operator-run supply, and gas-sponsored usage after the initial deposit.

Discussion insight: These posts did not disagree about whether more agents are coming. They disagreed about what the bottleneck is: permissions, portability, discovery, or settlement. But all four assumed the bottleneck lives outside the core prompt loop.

Comparison to prior day: August 4 focused on wallets, escrow, and spending controls. August 5 widened that lens into full end-to-end surfaces for internal reach, public discovery, and machine-paid usage.


2. What Frustrates People

Skill abundance creates a new routing and discovery problem

Severity: High. @mattpocockuk released (1,219 likes, 49 replies, 38,104 views, 747 bookmarks) a broad skills update, but the sharpest reply said the range was already “a bit overwhelming” and explicitly asked for a skills-helper that picks the right skill for the job. @rlaope responded (45 likes, 1 reply, 3,317 views, 60 bookmarks) by pitching one plugin as the operating layer that cuts through “cognitive load,” while @alex_prompter answered (8 likes, 2 replies, 3,443 views) with explicit triggers and “NOT FOR” boundaries inside SKILL.md. The coping pattern is curation, stronger boundaries, and wrapper layers above raw skill folders. This is worth building for because the friction appears precisely when a library becomes useful enough to grow.

Harnesses are still opaque enough that third parties have to explain them

Severity: High. @morganlinton called (505 likes, 3 replies, 129,314 views, 787 bookmarks) a harness-engineering guide his best read on the topic, and @swyx sent (82 likes, 10 replies, 8,623 views, 83 bookmarks) readers to an outside reconstruction of ChatGPT Work rather than an official product explainer. Replies under that post said harness writeups are becoming more useful than the official docs, while replies to @ren_hongyu's Muse Code launch (166 likes, 9 replies, 8,633 views, 8 bookmarks) complained that the product was hard to find on the public web and that the benchmark charts were not transparent enough. People cope by bookmarking teardown articles, comparing charts, and leaning on external explainers. This is worth building for because a hidden harness is hard to trust, compare, or reproduce.

Enterprise deployment still collapses into permissions, continuity, and access surfaces

Severity: High. @rileybrown said (71 likes, 8 replies, 7,636 views, 100 bookmarks) Vercel's internal agent already serves almost 1,000 people, but the discussion immediately moved to whether one agent should exist at all or whether teams need several agents with different permissions and computer access. The public ChatGPT Work reconstruction added another continuity complaint: cross-task context is mediated through product-layer services instead of a freely shared workspace. @Fetch_ai summarized (107 likes, 1 reply, 6,067 views) the same post-build burden as deployment, discovery, orchestration, user access, and payments. The workaround is to add more routing, more product-layer context, and more explicit surfaces after the model loop. This is worth building for because multiple independent posts described it as the real work after “building the agent.”

Agent finance and open inference still need a trust layer, not just a wallet

Severity: Medium. @ama_protocol framed (54 likes, 49 replies, 611 views) Amadeus around two missing pieces: a place to find and run agents, and a private-but-provable record of what they did with real money. @wardenprotocol launched (47 likes, 14 replies, 7,177 views) Halo with wallets, deposits, operator-run supply, and verifiable inference positioning, but both products are still selling the surrounding trust design more than mainstream usage proof. The coping pattern is escrow, budgets, provable records, and wallet-bounded access. This is worth building for, but today's evidence still looks early and infrastructure-heavy rather than broadly adopted.


3. What People Wish Existed

Skill choosers that sit above the skill library

Practical need. The clearest request came in a reply to @mattpocockuk's v1.2 launch (1,219 likes, 49 replies, 38,104 views, 747 bookmarks): the library is powerful, but broad enough to need a skills-helper that recommends what to run. @rlaope framed (45 likes, 1 reply, 3,317 views, 60 bookmarks) the same gap as plugin fatigue, and @alex_prompter answered (8 likes, 2 replies, 3,443 views) with stricter trigger and boundary rules. The missing product is a recommender and governor for skills, not just a bigger catalog. Opportunity: direct.

One front door for agent teams with explicit permissions and durable context

Practical and urgent need. @rileybrown made (71 likes, 8 replies, 7,636 views, 100 bookmarks) the “one GOD agent or many smaller agents” question explicit, and the strongest reply said the decision usually collapses to permissions. The ChatGPT Work teardown and @Fetch_ai's stack summary (107 likes, 1 reply, 6,067 views) both imply the same missing layer: a durable front door that can route work across agents, contexts, devices, and services without hiding continuity behind opaque product plumbing. Existing products partially address this, which makes the opportunity competitive rather than empty. Opportunity: competitive.

Benchmark surfaces that are public enough to audit, not just admire

Practical need. @morganlinton bookmarking (505 likes, 3 replies, 129,314 views, 787 bookmarks) an external harness guide, @swyx sharing (82 likes, 10 replies, 8,623 views, 83 bookmarks) an outside reconstruction of ChatGPT Work, and the replies challenging @ren_hongyu's Muse Code charts (166 likes, 9 replies, 8,633 views, 8 bookmarks) all point to the same desire: if a harness matters this much, people want public docs, reproducible evals, and context around what the numbers mean. This is partly a tooling problem and partly a reporting one. Opportunity: competitive.

Wallet-bounded execution with provable records for agent finance

Practical need, early market. @ama_protocol split (54 likes, 49 replies, 611 views) the product into a discovery hub and an embedded execution layer, both centered on visible action records. @wardenprotocol described (47 likes, 14 replies, 7,177 views) Halo as a wallet-driven inference market where agents can pay their own way from deposits. The missing product is broader than a wallet: identity, spend bounds, proof of action, and a place to find trustworthy services. Opportunity: direct.


4. Tools and Methods in Use

Tool Category Sentiment Strengths Limitations
mattpocock/skills Skill library / installer (+) Launch post (1,219 likes, 49 replies, 38,104 views, 747 bookmarks) added full docs, a Claude Code marketplace plugin, Codex support, and both managed and editable install paths A reply said the breadth is already overwhelming without a skill recommender
oh-my-hermes Hermes operating layer (+/-) Project post (45 likes, 1 reply, 3,317 views, 60 bookmarks) and the public site show planning, research, coding handoff, operations, and project memory across Hermes Desktop, CLI, and messenger Tied to Hermes; public evidence is still mostly builder-authored rather than peer-reviewed
Prime Agent Coding harness (+) Launch post (152 likes, 5 replies, 8,107 views, 51 bookmarks) described persistent IPython, programmatic subagents, agent messaging, and self-improving harness CRUD; the blog claims 95.5% ARC-AGI-3 with Opus 5 Early launch, and replies still asked about performance on cheaper or open models
Muse Code Coding harness (+/-) Launch post (166 likes, 9 replies, 8,633 views, 8 bookmarks) plus Meta's post emphasize persistent async background agents, replay-exact event logs, and bundled /plan, /grill, and /goal skills Replies challenged discoverability and benchmark transparency
ChatGPT Work Knowledge-work harness (+/-) Latent.Space reconstruction, shared via @swyx (82 likes, 10 replies, 8,623 views, 83 bookmarks), mapped cloud microVMs, browser use, plugins, artifacts, and product-managed cross-task context The strongest public explanation came from outsiders; continuity across tasks is still partly opaque
Fetch.ai / Agentverse / ASI:One Deployment / discovery stack (+/-) Stack summary (107 likes, 1 reply, 6,067 views) clearly maps build -> deploy/connect -> user reach -> trusted services Public site detail is thin and today's thread had little practitioner feedback
Halo AI inference marketplace (+/-) Launch post (47 likes, 14 replies, 7,177 views) and the launch blog describe operator-run model supply, agent wallets, USDC settlement, and gas-sponsored inference after deposit Still in public alpha, with crypto wallet onboarding and little proof of broad usage
Amadeus / AMA Agentic finance settlement layer (+/-) Product post (54 likes, 49 replies, 611 views) splits the product into a discovery hub and an embedded execution layer with private-but-provable records Evidence today came mostly from the builder's own framing, not independent operator/customer reports
Agent Arena Multi-agent evaluation / arena (+) Launch post (47 likes, 15 replies, 1,949 views) opened social-strategy games to the public, while the repo exposes judge, whisper mode, replay, and multi-provider orchestration Early metrics and methodology are still evolving

Overall satisfaction was highest when a tool exposed a clear operating surface: install paths, replay logs, explicit subagents, artifacts, or payment rails. The main workarounds were wrapper layers above raw skills, outside writeups for opaque harnesses, and wallet-bounded or product-managed context when direct cross-system continuity was too risky. The competitive dynamic is now less about one best model and more about who owns packaging, orchestration, eval visibility, and the post-build path to users.


5. What People Are Building

Project Who built it What it does Problem it solves Stack Stage Links
Skills for Real Engineers @mattpocockuk Large public skill library with docs, plugin distribution, and editable installs Makes reusable agent workflows easier to install, adapt, and share across coding agents Markdown skill packs, skills.sh installer, Claude Code plugin, Codex agents/openai.yaml support Shipped repo, tweet
Prime Agent @johannes_hage Self-improving coding and research harness with persistent REPL, subagents, and harness CRUD Extends coding agents to long-running work without freezing prompts, skills, or memory at design time TypeScript, persistent IPython kernel, rlm subagents, harness CRUD, autonomous mode Shipped blog, repo, tweet
Muse Code @ren_hongyu Terminal coding agent co-trained with Muse Spark 1.2 for large-repo engineering tasks Tries to make long-horizon coding more reliable through persistent helpers and replay-safe runtime state Muse Spark 1.2, async background agents, local event log, bundled /plan /grill /goal skills Beta blog, tweet
oh-my-hermes @rlaope Operating layer for Hermes that standardizes planning, research, handoff, operations, and memory Reduces plugin fatigue and adds clearer evidence boundaries above raw Hermes skills Python, Hermes Agent, managed skills, CLI/Desktop/Messenger workflows Shipped repo, site, tweet
Fetch.ai agent stack @Fetch_ai Connects build, deploy, discovery, user access, and trusted services into one agent path Helps teams move from building an agent to running a reachable service uAgents, Agentverse, ASI:One, Fetch Business Shipped site, tweet
Halo @wardenprotocol Peer-to-peer AI inference market where humans or agents can buy or serve inference with USDC Removes centralized API-gatekeeping and lets agents pay for inference from their own wallets Base, USDC, Halo vault, operator network, OpenAI-compatible local endpoint Alpha blog, repo, tweet
Agent Arena @sensho Public multi-agent game arena and eval surface with judges, replay, and multiple providers Gives builders a live environment for observing social strategy, preference, and deception behavior Next.js, TypeScript, Prisma, multi-provider model routing, judge/scoreboard, replay Beta repo, tweet

Matt Pocock's skills pack and oh-my-hermes represent the same builder pattern at different layers. One packages many reusable workflows for broad agent compatibility; the other wraps a specific agent with a stronger operating layer, evidence boundaries, and curated defaults. The repeated trigger is not model quality alone, but the management overhead that appears once skills start to accumulate.

Prime Agent and Muse Code formed the day's clearest harness race. Both emphasize long-running coding work, persistent helpers, and explicit runtime state, but Prime Agent pushed harder on self-modifying harness state while Muse Code stressed persistent async helpers and restart-safe logs. The shared build pattern is that harnesses are now being shipped as first-class products rather than quietly embedded around the model.

Fetch.ai, Halo, and Agent Arena broadened the build surface beyond coding. One organizes delivery and discovery, one organizes payment and supply, and one organizes public evaluation traffic. Together they suggest that more builders are moving from “can the agent act?” to “can users reach it, pay it, and compare it?”


6. New and Notable

Skills reached a true distribution milestone

@mattpocockuk reported (1,219 likes, 49 replies, 38,104 views, 747 bookmarks) 13.5 million installs for mattpocock/skills while adding docs, a Claude Code marketplace plugin, and Codex support. The notable shift is not just another repo update; it is that skills now have marketplace, installer, and documentation expectations closer to a product than a prompt pack.

Prime Agent pushed the self-improving harness story into the open

@johannes_hage launched (152 likes, 5 replies, 8,107 views, 51 bookmarks) Prime Agent, and @Dr_Singularity broadcast (343 likes, 21 replies, 20,024 views, 86 bookmarks) its 95.5% ARC-AGI-3 chart. What makes it notable is the combination of public architecture, public repo, and public benchmark rhetoric around a harness that edits its own prompts, memories, skills, and subagent definitions.

ChatGPT Work's context boundary got an unusually clear public explanation

@swyx shared (82 likes, 10 replies, 8,623 views, 83 bookmarks) Latent.Space's ChatGPT Work reconstruction. The notable detail was not just that Work uses browser tools and plugins; it was the explicit separation between per-task working directories and an OpenAI-managed layer for memory, projects, files, and personal context.

Multi-Agent Arena opened social-strategy evaluation to the public

@sensho opened (47 likes, 15 replies, 1,949 views) Multi-Agent Arena for free public play, then said in replies that early access had already passed a thousand matches per day and was approaching 1 billion tokens per day. That is notable because it turns “agents playing games” into a live evaluation surface with public traffic, anonymous preference votes, and behavioral observations such as model lying.

Grok shipped a supervision-oriented release instead of a model announcement

@cb_doge posted (81 likes, 40 replies, 11,209 views) Grok build v0.2.121 with previous-turn summaries in the dashboard, alphabetized Skills sections, session reattach without transcript replay, and fixes for background-task restarts and MCP image corruption. The release-notes image matters because it shows product work aimed at supervising agents already in flight rather than simply making another model claim.

Grok v0.2.121 release notes showing previous-turn summaries, grouped skills, parent-agent reminders, and session reattach without transcript replay


7. Where the Opportunities Are

[+++] Skill packaging, discovery, and governance layers — Matt Pocock's 13.5 million-install skills pack, the reply asking for a skills-helper, oh-my-hermes' attempt to sit above plugin sprawl, and the SKILL.md boundary template all point to the same gap: once skills become plentiful, somebody has to recommend, route, constrain, and maintain them. The evidence spans sections 1, 2, 3, and 5, which makes this the strongest near-term product opportunity in the dataset.

[+++] Permission-aware agent front doors — Riley Brown's Vercel interview made the “one GOD agent or many permissioned agents” tradeoff explicit; ChatGPT Work's public teardown exposed how much continuity is still mediated by product layers; and Fetch.ai framed deployment, discovery, orchestration, access, and payments as the real journey after the build. The consistent opportunity is a front door that can route work across agents, services, and contexts without losing control or continuity.

[++] Public harness observability and benchmark audit tools — Morgan Linton's bookmarking of a harness guide, Latent.Space's reverse engineering, Prime Agent's public architecture plus benchmark claims, and the pushback on Muse Code's charts show real demand for public explanation and auditable evaluation. The opportunity is meaningful, but competitive: many people can publish charts, while fewer can become the trusted surface for comparing them.

[+] Agent payment, inference, and trust rails — Amadeus and Halo both argued that wallets alone are not enough; discovery, provable records, operator supply, and spend bounds matter too. The evidence is concrete but still early, so the opportunity is emerging rather than fully validated.


8. Takeaways

  1. Skills are no longer just reusable prompts; they are packaged products with installers, docs, and governance problems. Matt Pocock's v1.2 release and the reply asking for a skills-helper showed both scale and the new discovery burden, while oh-my-hermes showed a parallel move toward wrapper layers above raw skills. (source)
  2. Harness engineering is now something people benchmark, reverse engineer, and publicly compare. The day's strongest evidence combined Morgan Linton's bookmarking of a harness guide, Latent.Space's reconstruction of ChatGPT Work, Prime Agent's open architecture, and Muse Code's contested charts. (source)
  3. The post-build surface has become the real engineering problem for many teams. Riley Brown's Vercel interview, Fetch.ai's build-to-service stack, and ChatGPT Work's product-managed continuity all pointed to permissions, context, discovery, and access as the hard parts after the agent exists. (source)
  4. Agent payment and inference rails are getting more concrete, but they still read as infrastructure bets rather than settled demand. Halo's public-alpha launch and Amadeus' hub/embed split both added specifics on wallets, records, and service discovery, yet neither thread showed broad user adoption as clearly as the packaging and harness themes did. (source)