Twitter AI Agent - 2026-09-12¶
1. What People Are Talking About¶
1.1 Harness engineering became the common language for serious agent work (🡕)¶
The strongest cluster no longer treated AI agents as "prompting" with a nicer wrapper. It treated them as an operational stack: concepts, loops, permissions, evals, context, recovery, and human judgment. The repeated move was to make agent behavior teachable, reviewable, and separable from the raw model.
@matthewcanham argued (93 likes, 11 replies, 6,976 views, 175 bookmarks) that every PM now needs at least 101-level literacy in tool use, context engineering, memory, harness engineering, evaluation, and sandboxing. The replies made the post more practical than inspirational: one reply said the fastest way to learn was to build a tiny tool call, not memorize vocabulary.
@iiiichigo_chan highlighted (69 likes, 11 replies, 8,961 views, 106 bookmarks) a free course on production-grade Claude Code harnesses covering Plan Mode, reusable skills, subagents, feedback loops, and recovery. The most useful replies turned that into an operator rule: a CLAUDE.md instruction only matters when it has a trigger, a true/false check, a stop condition, and evidence that it was followed.
@adiix_official turned (37 likes, 11 replies, 2,444 views, 20 bookmarks) Andrew Ng's AI-engineering framework into a two-page build checklist about context control, testing, security, and evaluation, while @DamiDefi amplified (140 likes, 10 replies, 11,456 views) the same shift from a different angle: code generation is getting cheaper, so product judgment and build-loop ownership are becoming scarcer than implementation labor. Even low-volume posts like @pauliusztin_ argued (5 likes, 3 replies, 182 views) that the agent loop itself is small and the hard part is the harness around it.

Discussion insight: the discourse moved one step past "AI engineering matters" into how to operationalize it. People kept asking for rules that survive natural-language drift, model changes, and long-running workflows.
Comparison to prior day: 2026-09-11 made AI engineering look like a measurable operating discipline. On 2026-09-12, it looked more like a teachable curriculum with courses, checklists, and named harness components.
1.2 Skills stopped looking like harmless snippets and started looking like software artifacts (🡕)¶
A second cluster focused on skills themselves: how to inspect them, measure whether they even trigger, and keep them from sprinting into implementation before the plan is stable. The shared premise was that skills are useful precisely because they change behavior, which also makes them risky.
@traversymedia launched (53 likes, 8 replies, 2,360 views, 44 bookmarks) SkillPass, an open-source directory where each skill gets a "passport" showing what it asks an agent to do before install. The linked public docs and repository make the trust model explicit: pinned commits, permission detection, validation findings, AI review, and hash verification at install time. Replies stated the pain plainly: blind-installing a skill can amount to quietly handing a shell your keys.
@stretchcloud reported (9 likes, 3 replies, 462 views) that claude plugin eval now runs each test case twice - with and without the plugin - so maintainers can see whether a skill actually changes outcomes. The most important first finding was behavioral, not syntactic: a plugin can validate cleanly yet still fail to trigger when a user phrases the request naturally instead of using the exact keyword from the manifest.
@Umesh__digital described (12 likes, 10 replies, 423 views) Matt Pocock's Grill-Me skill as a direct response to another failure mode: coding agents start building before the human has finished thinking. Its core heuristic - ask humans about decisions, ask the codebase about facts - fit neatly with the broader shift from one-shot prompting toward explicit design resolution.

Discussion insight: people did not assume skills are automatically helpful. They treated them as untrusted software artifacts that need permission visibility, trigger testing, and clearer decision boundaries.
Comparison to prior day: 2026-09-11 emphasized reviewable memory and policy gating. Today the same instinct moved outward into the skill ecosystem itself.
1.3 Agent commerce shifted from directories to trust rails and buyer scarcity (🡕)¶
The marketplace conversation was more concrete than yesterday, but the substance came from transaction design rather than generic future-of-agents rhetoric. The strongest posts all circled the same questions: what gets priced, how completion is verified, and whether enough real work exists on the buyer side.
@Abba__80 argued (152 likes, 176 replies, 976 views) that the most interesting TermiX metric was not the headline $19.68M settled, but the implied roughly $53 average job size across 370,319 jobs. The attached dashboard also showed 431,692 total agents and $393,618 in protocol revenue, which made the thesis legible: agent commerce may emerge through many small machine-to-machine transactions rather than a few giant contracts.
@dee_e6 framed (58 likes, 63 replies, 321 views) AACP and agent.family as the layer that links identity, bidding, escrow, delivery verification, dispute handling, and repeatable work history. The image made that more concrete by showing a public reputation score, completed-job count, and end-to-end job flow. @dezydank supplied (73 likes, 58 replies, 734 views) the promotional version - funds stay locked until work is independently checked - but replies immediately attacked the real boundary by asking who performs that verification and under what rules.
@Arifx0001 surfaced (81 likes, 84 replies, 10,486 views) and @ZeoWeb3 supplied (77 likes, 75 replies) the sharpest demand-side numbers in the set: a live marketplace surface with roughly 584 sellers or services versus 224 buyers or requests, plus 0 active bounties in one screenshot. @icefrog_sol distilled (14 likes, 15 replies, 93 views) the buyer-side need to a simple rule: budget, deadline, and pass/fail test should be defined before the work begins.



Discussion insight: even supporters kept dragging the conversation back to acceptance tests, recourse, and demand generation. The interesting bottleneck was not listing another agent; it was bringing verified work to the marketplace.
Comparison to prior day: 2026-09-11 said market mechanisms were getting more concrete while demand proof stayed thin. On 2026-09-12, buyer scarcity and verification became the main lens rather than a side note.
1.4 Runtime choice widened fast: local small models, managed sessions, reusable browser skills, and cloud/local handoff (🡕)¶
A fourth cluster showed the stack spreading in multiple directions at once. Instead of one dominant runtime pattern, people compared local small models, managed agent infrastructure, browser skill factories, cloud-to-local workspaces, and voice-agent primitives.
@aibytekat argued (56 likes, 20 replies, 4,031 views) that MiniCPM5-2B makes practical local agent workflows plausible on ordinary hardware. The OpenBMB release materials support that framing directly: the model is pitched for local assistants, coding agents, tool use, and other resource-constrained agent tasks. Replies made the boundary clearer by asking about hardware floors and whether structured tool-calling fidelity survives quantization.
@beamnxw explained (25 likes, 18 replies, 520 views, 11 bookmarks) the new OpenAI Agents API as a split between durable sessions, an execution box, and a vault that keeps secrets out of that box. The official docs confirm the same pattern: OpenAI can manage orchestration, context compaction, recovery, and optional subagents while the environment can be OpenAI-hosted, self-hosted, or absent entirely. Replies zeroed in on call-time credential injection as the real security improvement.
@DivyanshT91162 surfaced (6 likes, 1 reply, 284 views) Microsoft's Webwright as a browser-agent approach where solved tasks become reusable verified code skills instead of one-off browsing traces. @euboid praised (25 likes, 10 replies, 2,088 views, 22 bookmarks) Conductor for multi-agent cloud workspaces, local sync, and handoff quality, only for replies to note the remaining rough edges: Mac-only constraints and iOS still being TestFlight-only. On the voice side, @minchoi showed (60 likes, 10 replies, 10,064 views, 31 bookmarks) full-duplex GPT-Live-1 notable because the improvement was concrete enough to price at $0.05 per minute.

Discussion insight: tool comparisons were increasingly about durability, reusability, and boundary management - where state lives, how work survives sleep, how secrets stay out of the sandbox, and how outputs become reusable artifacts.
Comparison to prior day: 2026-09-11 stayed centered on hooks, traces, and memory discipline. Today the conversation widened into runtime selection and deployment shape.
2. What Frustrates People¶
Skills are still too easy to trust blindly and too hard to evaluate¶
Severity: High. The skill discussion was full of people discovering that a skill can be dangerous, invisible, or both. @traversymedia framed (53 likes, 8 replies, 2,360 views, 44 bookmarks) the problem as blind trust: users often install agent skills without seeing what they will actually ask the agent to do. @stretchcloud reported (9 likes, 3 replies, 462 views) the adjacent frustration: even a well-formed plugin may do nothing because it fails to trigger on natural language, a gap that ordinary validation does not catch.
The workaround pattern is revealing. SkillPass adds a permission-and-passport layer before install, while Grill-Me asks clarifying questions before the agent writes code. That means the current coping strategy is not more model cleverness; it is more preflight inspection and more structured planning.
Worth building for? Yes. The pain is concrete, repeated, and close to the workflow edge. A product that combines permission visibility, trigger testing, and spec clarification would meet an obvious need.
Agent marketplaces still lack buyer-side proof and enough paid work¶
Severity: High, though confidence should stay moderate because many posts were promotional. @ZeoWeb3 posted (77 likes, 75 replies) the sharpest imbalance in the set: roughly 584 services listed versus 224 requests waiting, with 0 active bounties. @dee_e6 added (58 likes, 63 replies, 321 views) the deeper diagnosis: demos are weaker than work history, but work history still needs verification, settlement, and dispute handling to be trustworthy.
@dezydank framed (73 likes, 58 replies, 734 views) and @icefrog_sol showed (14 likes, 15 replies, 93 views) the same boundary from opposite directions. Escrow sounds reassuring until buyers ask who verifies completion; buyer-side demand sounds easy until someone has to define budget, deadline, and pass/fail criteria precisely enough for a machine-to-machine transaction.
Worth building for? Yes, but only if the product solves proof and demand together. More listings alone did not look like the bottleneck.
The runtime layer still leaks friction at the boundaries¶
Severity: Medium. The runtime discussion made clear that useful agent work still breaks at the seams: where secrets live, how cloud work hands back to local tools, how long a workspace survives, and what class of hardware can run a local model well enough to be worth the trouble. @beamnxw put (25 likes, 18 replies, 520 views, 11 bookmarks) the security version bluntly: a sandbox that can execute code should not also permanently hold the credential. @euboid praised (25 likes, 10 replies, 2,088 views, 22 bookmarks) Conductor's handoff, but replies and docs still surfaced Mac-only limitations, one-way sync, and workspace lifetime boundaries.
The small-model and voice-agent threads showed the same theme in different form. @aibytekat argued (56 likes, 20 replies, 4,031 views) for local placement, but got immediate questions about whether a 2B model stays reliable enough on commodity devices, while replies to @minchoi showed (60 likes, 10 replies, 10,064 views, 31 bookmarks) that smoother voice turn-taking only matters if downstream task completion is still trustworthy.
Worth building for? Yes. These are not abstract complaints; they are deployment constraints that shape which agent workflows people can actually keep.
3. What People Wish Existed¶
Verified skill distribution that also measures real trigger behavior¶
The skill cluster points to a missing layer above raw installation. @traversymedia argued (53 likes, 8 replies, 2,360 views, 44 bookmarks), @stretchcloud reported (9 likes, 3 replies, 462 views), and @Umesh__digital described (12 likes, 10 replies, 423 views) together suggest people want one place to inspect permissions, understand a skill's behavior, test whether it triggers on natural requests, and clarify the plan before execution starts.
Why now: skills are spreading quickly across coding-agent ecosystems, but the trust chain is still fragmented between install UX, evaluation, and planning.
Buyer-side control planes for agent work¶
The stronger ask in the marketplace cluster was not another directory. It was infrastructure that makes a job purchasable: scoped requirements, budget, deadline, acceptance tests, escrow, reputation, dispute handling, and evidence that the work was actually done. @icefrog_sol argued (14 likes, 15 replies, 93 views), @dee_e6 framed (58 likes, 63 replies, 321 views), and @ZeoWeb3 showed (77 likes, 75 replies) all pointed toward the same missing transaction layer.
Why now: public supply is visible already. What still looks scarce is buyer confidence and buyer-generated work.
Hybrid runtime orchestration across local models, managed sessions, and browser skills¶
The runtime discussion implied a wish for an orchestration layer that can place the right work on the right substrate: small local models for cheap private tasks, durable managed sessions for long runs, browser skill libraries for repeated flows, and cloud/local handoff that does not force teams to choose one environment forever. @aibytekat argued (56 likes, 20 replies, 4,031 views), @beamnxw explained (25 likes, 18 replies, 520 views, 11 bookmarks), @DivyanshT91162 surfaced (6 likes, 1 reply, 284 views), and @euboid praised (25 likes, 10 replies, 2,088 views, 22 bookmarks) all hint at pieces of that future.
Why now: teams are already mixing runtimes. The missing product is the control surface that makes those mixes reliable.
Eval, sandbox, and accountability infrastructure for long-horizon agents¶
The research cluster kept converging on a similar wish: cleaner evaluation boundaries, reusable environments, and external accountability when the stakes are high. @marfinxx highlighted (31 likes, 7 replies, 1,351 views, 29 bookmarks), @mirku21 highlighted (10 likes, 3 replies, 277 views, 7 bookmarks), and @gurtej__gill_ argued (19 likes, 2 replies, 665 views, 21 bookmarks) for evaluation design and environment decoupling, while @JoshAEngels made (164 likes, 8 replies, 2,929 views, 19 bookmarks) the accountability case concrete by joining METR.
Why now: once agents run longer and act more independently, evaluation and governance stop being research afterthoughts and become core infrastructure.
4. Tools and Methods in Use¶
Skill passports and permission review before install¶
The clearest method shift was to treat skills like software packages instead of helper text. @traversymedia used (53 likes, 8 replies, 2,360 views, 44 bookmarks) SkillPass to surface pinned commits, permissions, validation findings, and AI-readable summaries before install. The method matters because it puts inspection ahead of execution.
Interview-first planning before code generation¶
@Umesh__digital described (12 likes, 10 replies, 423 views) Grill-Me's most useful technique plainly: ask humans about decisions, ask the codebase about facts. That turns planning into a structured interview rather than letting the agent silently choose architecture, framework, and scope.
With/without delta evaluation for skills¶
@stretchcloud surfaced (9 likes, 3 replies, 462 views) a simple but important evaluation method: run the same test cases with and without the plugin and measure the delta. This is stronger than schema validation because it tests whether the skill changes behavior under natural phrasing.
Durable sessions with secret isolation¶
@beamnxw described (25 likes, 18 replies, 520 views, 11 bookmarks) the OpenAI Agents API as a session-plus-sandbox-plus-vault pattern. The relevant method is not merely “managed agents,” but separating orchestration state from code execution and keeping credentials outside the agent's box until call time.
Browser tasks captured as rerunnable code skills¶
@DivyanshT91162 highlighted (6 likes, 1 reply, 284 views) Webwright's method: treat the browser as an environment a coding agent can script, not as the persistent state machine. Solved tasks leave behind runnable Playwright code that can later be reused as a verified skill instead of re-derived token by token.
Cloud/local tandem workspaces and small-model placement¶
Two different methods surfaced here. @euboid emphasized (25 likes, 10 replies, 2,088 views, 22 bookmarks) shared cloud workspaces with local sync or SSH handoff, while @aibytekat argued (56 likes, 20 replies, 4,031 views) for pushing suitable agent tasks onto a 2B local model. Together they show a placement mindset: choose the environment that matches the task's privacy, latency, and cost profile.
Hidden evaluation and decoupled rollouts for long-horizon agents¶
The research items exposed a more specialized but important method shift. @marfinxx emphasized (31 likes, 7 replies, 1,351 views, 29 bookmarks) hidden consistent evaluation and asynchronous worker pools in AIRA-2; @mirku21 highlighted (10 likes, 3 replies, 277 views, 7 bookmarks) ProRL Agent's rollout-as-a-service split; and @gurtej__gill_ focused (19 likes, 2 replies, 665 views, 21 bookmarks) on Orchard's harness-agnostic environment layer. The common method is to separate evaluation and execution concerns so longer agent runs do not collapse under their own infrastructure.
5. What People Are Building¶
| Project | Who built it | What it does | Problem it solves | Stage | Links |
|---|---|---|---|---|---|
| SkillPass | Brad Traversy | Validation-first skill directory and CLI with immutable Skill Passports, permission detection, and hash-verified installs | Lets users inspect what a skill will do before trusting it on their machine | Public launch / open source | tweet, docs, repo |
| Claude plugin eval | Anthropic / Claude Code, surfaced by @stretchcloud | Runs skills against test cases with and without the plugin, then compares the delta with graders like tool_used and tool_order |
Measures whether a skill actually triggers and improves behavior, not just whether the manifest validates | Shipped feature | tweet, announcement |
| Grill-Me | Matt Pocock, surfaced by @Umesh__digital | Turns a coding agent into a technical interviewer that asks clarifying questions and recommends answers before implementation | Prevents agents from choosing architecture and scope before the plan is fully resolved | Open-source skill | tweet |
| TermiX / agent.family / AACP | TermiX ecosystem, surfaced by @Abba__80, @Arifx0001, and @dee_e6 | Marketplace and protocol stack for agent identity, bidding, escrow, verification, settlement, and reputation | Makes agent work easier to price, verify, and buy again | Live marketplace plus protocol build-out | market tweet, site, whitepaper |
| MiniCPM5-2B | OpenBMB | Compact 2B model positioned for local assistants, coding agents, tool use, and on-device deployment, with open training data and agent skills | Makes local or offline agent workflows more plausible on everyday hardware | Newly released open model | tweet, repo |
| OpenAI Agents API | OpenAI, surfaced by @beamnxw | Managed Codex harness with durable sessions, tool calling, sandbox choices, recovery, and optional subagents | Offloads orchestration and state management for long-running agents | Shipped API | tweet, docs |
| Webwright | Microsoft | Browser-agent framework where a coding model writes Playwright scripts and distills successful runs into reusable skills | Converts brittle one-off browsing runs into rerunnable code artifacts | Open source | tweet, repo |
| Conductor Cloud | Conductor, surfaced by @euboid | Shared cloud workspaces with isolated sandboxes, SSH or local-sync access, and multiplayer handoff | Improves cloud/local developer experience for agent workspaces | Product / shipped docs | tweet, docs |
| AIRA-2 | Meta FAIR, surfaced by @marfinxx | Autonomous research agent with hidden consistent evaluation, stateful ReAct loops, and asynchronous multi-GPU search | Tackles long-horizon evaluation and compute bottlenecks in research agents | Research paper | tweet, paper |
| Orchard | Microsoft Research, surfaced by @gurtej__gill_ | Open-source agentic modeling framework with a harness-agnostic, Kubernetes-native environment layer | Standardizes environments and trajectory reuse across agent tasks and harnesses | Research framework / open source | tweet, paper |
| ProRL Agent | NVIDIA, surfaced by @mirku21 | Rollout-as-a-service infrastructure that decouples multi-turn agent execution from RL training | Reduces GPU idle time and makes rollout infrastructure more portable across sandboxes | Research infrastructure / paper | tweet, paper |
The build pattern was unusually consistent across these projects. Builders were not chasing one monolithic “agent platform.” They were carving the stack into trust layers, evaluation surfaces, runtime choices, reusable skills, and environment infrastructure.
6. New and Notable¶
6.1 Usage data made the shift from chat windows to agent fleets easier to quantify¶
@beamnxw summarized (11 likes, 5 replies, 233 views, 6 bookmarks) a Codex paper arguing that agent adoption grew 5x in six months, 26.6% of users now use reusable skills, and more than 10% manage three or more agents concurrently in a typical week. The most interesting reply did not contest adoption; it asked how many of those agents still have an owner carefully reviewing their work. That made the paper notable for the same reason it was useful: it quantified a workflow shift without resolving the control question.
6.2 The small-model local-agent claim crossed from theory into release framing¶
MiniCPM5-2B was not just another open model drop. @aibytekat argued (56 likes, 20 replies, 4,031 views) and the OpenBMB release materials both framed it explicitly as a local coding-agent and tool-use model. That matters because the debate immediately moved to practical constraints - hardware thresholds, quantization, and structured output reliability - instead of whether local agent use is conceptually interesting.
6.3 Voice agents cleared a smaller but important UX breakpoint¶
@minchoi made (60 likes, 10 replies, 10,064 views, 31 bookmarks) GPT-Live-1 notable by reducing the improvement to two operational facts: full-duplex interruption handling and a price of $0.05 per minute. Replies sharpened why that matters. The hard problem is not just lower latency; it is deciding in real time whether an interruption is a correction, a backchannel, or a new instruction.
6.4 Long-horizon research focus stayed on evaluation and environment design¶
The most substantive research items all argued that the bottleneck is infrastructure, not a magic prompt. @marfinxx highlighted (31 likes, 7 replies, 1,351 views, 29 bookmarks) Meta's AIRA-2 paper, which claims 83.1 percentile rank on MLE-bench-30 after 72 hours by combining hidden consistent evaluation, stateful ReAct loops, and asynchronous multi-GPU search. @mirku21 emphasized (10 likes, 3 replies, 277 views, 7 bookmarks) ProRL Agent's rollout-as-a-service split between trainers and execution infrastructure, while @gurtej__gill_ focused (19 likes, 2 replies, 665 views, 21 bookmarks) on Orchard's harness-agnostic environment layer.


What made these posts notable was the shared diagnosis. Longer-horizon agent progress depended on cleaner evaluation splits, reusable sandbox layers, and more efficient search or rollout infrastructure - not just larger models.
6.5 External evaluation remained a live governance story¶
@JoshAEngels announced (164 likes, 8 replies, 2,929 views, 19 bookmarks) that he left Google DeepMind's AGI safety team to join METR because he thinks the stakes are now high enough that more outside accountability is needed. Replies challenged definitions of “aligned” and “safe,” but that disagreement is part of why the post mattered: evaluation is no longer only a product or research concern. It is also becoming a public governance surface.
7. Where the Opportunities Are¶
[+++] Verified skill trust layers - This looked like the cleanest near-term opportunity because the pain appeared in multiple adjacent forms at once: unsafe installs, invisible permissions, weak trigger behavior, and premature execution. @traversymedia framed (53 likes, 8 replies, 2,360 views, 44 bookmarks), @stretchcloud reported (9 likes, 3 replies, 462 views), and @Umesh__digital described (12 likes, 10 replies, 423 views) all pointed toward the same missing layer: inspect the skill, test the trigger, clarify the plan, then install or run it.
[+++] Buyer-side control planes for agent commerce - The marketplace cluster did not need more agent listings nearly as much as it needed scoped jobs, acceptance tests, escrow, dispute handling, reputation, and demand creation. @Abba__80 argued (152 likes, 176 replies, 976 views), @ZeoWeb3 posted (77 likes, 75 replies), @dee_e6 framed (58 likes, 63 replies, 321 views), and @icefrog_sol argued (14 likes, 15 replies, 93 views) all suggest that “bring verified work” is the real unsolved problem.
[++] Heterogeneous runtime orchestration and handoff - People are clearly not converging on one runtime. They are mixing local small models, managed sessions, browser skill frameworks, cloud sandboxes, and voice layers. @aibytekat argued (56 likes, 20 replies, 4,031 views), @beamnxw explained (25 likes, 18 replies, 520 views, 11 bookmarks), @euboid praised (25 likes, 10 replies, 2,088 views, 22 bookmarks), and @minchoi showed (60 likes, 10 replies, 10,064 views, 31 bookmarks) point to a product space around workload placement, secret handling, session portability, and cloud/local continuity.
[++] Reusable task compilers for browser and coding workflows - Webwright and Grill-Me suggest a broader pattern: one class of tools helps agents think through the problem before they act; another captures solved work as code or skills that can be rerun without repeating the full reasoning loop. @DivyanshT91162 surfaced (6 likes, 1 reply, 284 views), @Umesh__digital described (12 likes, 10 replies, 423 views), and @pauliusztin_ argued (5 likes, 3 replies, 182 views) all support products that turn agent work into reusable, inspectable operational assets.
[+] Evaluation and sandbox infrastructure for long-horizon agents - This is less immediate for everyday buyers, but the technical signal was strong. @marfinxx highlighted (31 likes, 7 replies, 1,351 views, 29 bookmarks), @mirku21 highlighted (10 likes, 3 replies, 277 views, 7 bookmarks), @gurtej__gill_ argued (19 likes, 2 replies, 665 views, 21 bookmarks), and @JoshAEngels announced (164 likes, 8 replies, 2,929 views, 19 bookmarks) all point toward the same frontier: better eval boundaries, reusable environments, and stronger external accountability as agents run longer and matter more.
8. Takeaways¶
- Harness engineering is now the default language for serious agent work. The strongest posts framed progress in terms of loops, permissions, evals, recovery, and product judgment rather than prompt phrasing alone. (source, 93 likes, 11 replies, 6,976 views, 175 bookmarks)
- Skills are being treated more like untrusted software than harmless snippets. Permission passports, trigger evals, and interview-first planning all point to the same shift: skill usefulness has to be earned and inspected. (source, 53 likes, 8 replies, 2,360 views, 44 bookmarks)
- Agent commerce looked more legible, but buyer scarcity and verification still dominate the risk surface. Public metrics and marketplace screenshots made the category easier to reason about, yet the buyer side still appears thinner than the seller side and acceptance criteria remain underbuilt. (source, 77 likes, 75 replies)
- Runtime choice diversified quickly. The same day included local 2B agent models, managed session APIs, browser skill factories, cloud/local workspaces, and full-duplex voice primitives, which suggests the stack is fragmenting by task and operating constraint. (source, 25 likes, 18 replies, 520 views, 11 bookmarks)
- Research and governance signals converged on the same lesson: infrastructure matters as much as the model. Hidden evaluation, rollout services, reusable environment layers, and outside eval organizations all point to a future where trust depends on better harnesses and better accountability, not just better weights. (source, 31 likes, 7 replies, 1,351 views, 29 bookmarks)