Skip to content

Twitter AI - 2026-09-21

1. What People Are Talking About

1.1 MiMo-V2.6 reset the open-weight ceiling for multimodal agent work πŸ‘•

MiMo was the feed's clearest breakout. Roughly 10 MiMo/Xiaomi posts appeared on 2026-09-21 versus 2 the prior day, and the conversation mixed launch excitement with deployer math: active parameters, per-task cost, repo size, and whether the new weights were practical for local or self-hosted use.

@XiaomiMiMo announced (1,819 likes, 99 replies, 72,384 views, 318 bookmarks) MiMo-V2.6 Pro and Flash as fully open models with agent, multimodal, and 3D/computer-use ambitions, saying Pro hits 46 on the Artificial Analysis Intelligence Index and performs on par with Claude Opus 5 and GPT-5.6 Sol on many agent benchmarks. The public release note sharpens the thesis behind the launch: Xiaomi is framing the model as scaled reinforcement learning for self-improvement, not just a static checkpoint drop, and says the series ships with weights, a technical report, RL environments, training code, and unchanged API pricing versus V2.5.

Benchmark table from Xiaomi comparing MiMo-V2.6 Pro and Flash with MiMo-V2.5 Pro, Claude Opus 5, GPT-5.6 Sol, and Fable 5 across coding, general-agent, cyber, and visual tasks

@LuminaBench flagged (106 likes, 6 replies, 3,192 views) the deployer-side consequence: MiMo-V2.6-Pro enters the Artificial Analysis frontier at roughly $0.13 per task while a quoted benchmark note describes it as a 1.02T-parameter MoE with 42B active parameters. That made the model interesting not just as a leaderboard event but as a candidate for real routing-table experiments against other open models.

Artificial Analysis chart showing MiMo-V2.6-Pro entering the top open-weight tier at 46 while sitting near the cost-per-task Pareto frontier

@TeksEdge added (41 likes, 1,711 views) the hardware nuance that kept the thread grounded: Flash is a 309B-total / 15B-active model with a roughly 178 GB official repo, while Pro is around 573 GB, both with 1M context, multimodal IO, tool use, and documented SGLang/vLLM deployment. That post shifted the conversation from "top open model" to "what memory class do I need to run this locally, and which of the two variants is realistic first?"

Dark benchmark table from a MiMo deployment thread showing Pro and Flash side by side across DeepSWE, AutomationBench, Toolathlon, Terminal Bench, cyber, and visual tasks

Discussion insight: The most thoughtful reactions were not cheering the index number in isolation. They were about whether MiMo deserves a live A/B slot in production routing tables, because open weights plus strong benchmark scores plus unchanged pricing is finally close enough to make real workload comparisons worth the effort.

Comparison to prior day: Compared with 2026-09-20, when open-weight attention centered on Qwen-Image-2.1 and gateway share, 2026-09-21 consolidated the theme around one flagship launch and pushed the conversation deeper into cost, storage, and deployment constraints.

1.2 Grok 4.7 spread quickly, but as a price-adjusted engineering debate rather than a clean frontier win πŸ‘•

Grok 4.7 mentions jumped from 1 post on 2026-09-20 to 21 posts on 2026-09-21. The attention was real, but the tone was mixed: people liked the lower prices and some benchmark wins, while immediately questioning benchmark provenance, task fit, and whether the new model actually improves long-running engineering work.

@cb_doge claimed (259 likes, 45 replies, 17,697 views) that Grok 4.7 beat GPT-6 Astra on GDPval and AA Briefcase, using a chart that showed 1,695 vs 1,542 on professional knowledge work and 1,657 vs 1,569 on multi-hour office work. The distinctive part of the discussion came in replies, where people immediately asked about sample size, task mix, independence of the runs, and what happens after cost and retry rates are included.

Benchmark card showing Grok 4.7 ahead of GPT-6 Astra on GDPval and AA Briefcase Elo-style work benchmarks

@cognition rolled it into Devin (93 likes, 11 replies, 6,908 views) the same day, reporting a 59.4% FrontierCode 1.1 Extended score and saying the model performs very well on hard backend engineering tasks. Just as important, Cognition's own follow-up reply said Grok 4.7 still trails Grok 4.6 in aggregate because it over-scopes some tasks, which turned the post into one of the day's clearest examples of tool vendors refusing to oversell a flashy release.

FrontierCode 1.1 score-versus-cost chart showing Grok 4.7 below GPT-6 Astra on peak score but positioned at a lower average rollout cost than several other frontier coding models

@ChrisGPT summarized (29 likes, 9 replies, 5,184 views) the tradeoff in a way many smaller builders care about: Grok 4.7 sits at $2 / $6 per million input/output tokens, well below GPT-5.6 Sol Max and Fable 5.1 Max, while still posting a 71.0 DeepSWE score. The same comparison also showed why the celebration stayed qualified - Terminal Bench 4.0 remained well behind Fable 5.1 Max, so the model reads as attractive for some engineering workloads without yet closing every gap.

Comparison table showing Grok 4.7 pricing and benchmark results against Grok 4.6, GPT-5.6 Sol Max, and Fable 5.1 Max, including a weaker Terminal-Bench row

Discussion insight: The feed did not treat Grok 4.7 as a one-number frontier coronation. It treated it as a routing question: where is the new price/performance envelope genuinely better, and where do overscoping and terminal-task weaknesses erase the savings?

Comparison to prior day: Compared with the almost nonexistent 2026-09-20 Grok 4.7 chatter, 2026-09-21 turned it into a full-stack debate involving benchmark screenshots, live coding-tool integration, and explicit price tables.

1.3 Generation and control kept separating into different layers πŸ‘•

Another strong thread was architectural rather than model-specific. Roughly 38 Jev/classifier/control-layer posts and well over 100 benchmark/evaluation posts kept circling the same idea: let large models generate when needed, but peel routing, scoring, verification, and harness improvement into cheaper or more verifiable layers.

@ClementDelangue argued (84 likes, 13 replies, 6,554 views, 19 bookmarks) that frontier LLM APIs are overkill, slow, expensive, and hard to control for much of real-world AI, quoting a Jev thread that framed zero-shot classification as an old idea worth rescuing. @0xwhrrari made the architecture explicit (25 likes, 8 replies, 524 views, 21 bookmarks): routing, scoring, and verification become a fast decision layer, while generation stays elsewhere.

@0xClodex turned that into a playbook (8 likes, 1 reply, 216 views, 10 bookmarks) with a checklist for dynamic menus, confidence gates, fallback ladders, and the rule that deterministic limits and permissions should stay in code instead of being handed to the model. The image was useful because it showed the emerging split in one line: "LLM generates -> Jev decides -> code executes -> human reviews uncertainty."

Infographic listing 14 Jev tips, including dynamic menus, batched questions, confidence gating, fallback ladders, and the split between generation, decision, code execution, and human review

@whitecircle launched (171 likes, 66 replies, 15,372 views, 120 bookmarks) Halo as the same mindset applied one layer lower in the stack: a Hugging Face-native post-training framework with asynchronous RL and expert/context/tensor parallelism that claims up to 2.8x stock TRL throughput without abandoning the standard checkpoint format. @dair_ai highlighted (52 likes, 14 replies, 3,731 views, 37 bookmarks) ModularRSI for harness improvement, and the public repo says it evolves five harness modules separately and improves Terminal-Bench 2.0 accuracy from 47.57% to 52.43% on a benchmark-disjoint evolution pool.

Paper abstract screenshot for ModularRSI showing the benchmark-disjoint, five-module approach to recursive harness self-improvement

@OfficialLoganK put the management version bluntly (37 likes, 8 replies, 2,365 views): if you are building with AI, you should spend more than 25% of your time on benchmarks and on getting labs to care about them. That line fit the rest of the theme: people still want bigger models, but the higher-signal talk was increasingly about typed decisions, reproducible evaluations, and harness behavior.

Discussion insight: The interesting shift was not "agents are hot." It was that the community kept drawing boundaries around where uncertainty is acceptable and where it is not: keep rules in code, gate on confidence, validate outcomes, version the harness, and use cheaper models for repeated forks.

Comparison to prior day: On 2026-09-20, Jev talk was still mostly about where a decision-native model belongs. On 2026-09-21, the conversation broadened into concrete training frameworks, harness self-improvement methods, and explicit benchmark budgets.

1.4 Physical-AI posts stayed fixed on data engines, cheap capture, and trustable world-state APIs πŸ‘’

Physical-AI/data-engine discussion stayed almost flat day over day, with about 25 relevant posts on 2026-09-21 versus 24 on 2026-09-20. The difference was that more of the surviving posts were concrete about collection interfaces, verification loops, and API surfaces instead of speaking generically about "more data."

@itsbac0201 argued (12 likes, 9 replies) that Axis Robotics is interesting because it lets anyone teleoperate simulated robots in a browser and turns those sessions into quality-checked trajectories, with August figures of 200k+ registered users, 4.7m+ trajectories, 13 robot embodiments, and 160k+ downloads for Sim Dataset V1. The public AXIS project page makes the workflow more concrete: browser teleoperation feeds backend validation, smoothing, augmentation, and VLA training across a 207-task / 50K+ trajectory dataset.

@HuuHoang88 made the cheap-capture argument directly (22 likes, 24 replies, 205 views): spatial data is expensive when it requires special sensors and fleets, and Vangrid's wager is that the sensor is already in your pocket. The public API overview shows the same system being presented as an early-access product surface for spatial queries, data ingestion, and live ground-truth streaming.

@KaiBGR added the hardest caveat (14 likes, 16 replies, 1,206 views): Vangrid's verification design creates a real tradeoff between trust and speed, because multiple independent scans and corroboration improve confidence but add delay exactly where drones and robots want millisecond answers. That tradeoff was easiest to grasp in the attached diagram.

Diagram showing multiple device streams feeding Vangrid verification and aggregation steps before producing a high-trust spatial feed, with latency increasing as corroboration rises

Discussion insight: The disagreement was no longer whether physical AI needs better data. It was about what kind of data rail can survive contact with robotics requirements: cheap enough to collect, queryable enough to integrate, and fast enough that trust building does not make the signal useless.

Comparison to prior day: Compared with 2026-09-20's similar data-first framing, 2026-09-21 added more specific numbers around browser teleoperation and a clearer latency-versus-verification argument for real-time spatial feeds.


2. What Frustrates People

Benchmark wins still fail the workflow and P&L test

The sharpest frustration was that impressive charts still do not answer the operating question. @MatznerJon asked (26 likes, 9 replies, 4,432 views) whether any shiny new model, agent stack, or memory system actually improves contribution margin at the constraint or on-time task completion away from it; @OfficialLoganK said (37 likes, 8 replies, 2,365 views) teams should spend more than 25% of their time building benchmarks and trying to get labs to care about them; and @DataChaz boosted (22 likes, 1 reply, 1,178 views) a Codos claim that AI is crushing benchmarks while enterprises still struggle to see P&L impact. Even positive Grok 4.7 posts ran into the same wall: @cognition said the model can be strong on hard backend work while still trailing Grok 4.6 overall because it over-scopes some tasks.

Severity: High. Teams are coping with custom evals, tool-integrated benchmarks, and workflow-level success metrics, but the repeated demand for business-grounded evidence makes this clearly worth building for.

Tiny decisions are still too expensive when handled by general LLM calls

A second frustration was cost mismatch. @ClementDelangue argued (84 likes, 13 replies, 6,554 views, 19 bookmarks) that frontier LLM APIs are too overkill, slow, expensive, and hard to control for many real-world uses, while @0xwhrrari described (25 likes, 8 replies, 524 views, 21 bookmarks) Jev as the missing decision layer for routing, scoring, and verification. @0xClodex spelled out (8 likes, 1 reply, 216 views, 10 bookmarks) the coping pattern: dynamic menus, confidence gates, fallback ladders, and hard rules left in deterministic code rather than paid for over and over through text generation.

Severity: High. The workaround pattern is already stable enough to look like a product category: generation stays with a larger model, and repeated yes/no or ranking forks get pushed into cheaper typed decision systems.

Training and harness tooling still fractures once teams outgrow prototype scale

The feed also showed a persistent tooling cliff between early experimentation and serious post-training. @whitecircle launched (171 likes, 66 replies, 15,372 views, 120 bookmarks) Halo precisely around that pain, promising Hugging Face-native distributed training and claiming 2.3-2.8x stock TRL throughput without forcing teams into a new checkpoint format. @dair_ai surfaced (52 likes, 14 replies, 3,731 views, 37 bookmarks) the same fragility one layer higher: ModularRSI is needed because monolithic harness changes easily turn into benchmark-fitting instead of reusable improvement, and replies immediately asked for versioned traces, clean rollback, and better handling at module boundaries.

Severity: Medium-High. Builders are coping with HF-native parallelism, modular harnesses, and more validation gates, but the volume of infrastructure work around these topics suggests the prototype-to-scale transition is still painful and under-tooled.

Physical-AI data is still expensive to collect and slower to trust than robots want

Physical-AI posters kept describing the same upstream bottleneck in different forms. @itsbac0201 said (12 likes, 9 replies) Axis tries to scale manipulation data through browser teleoperation and quality-checked uploads; @HuuHoang88 argued (22 likes, 24 replies, 205 views) that Vangrid cuts collection cost by using phones as sensors; and @KaiBGR warned (14 likes, 16 replies, 1,206 views) that high-trust verification adds delay exactly where autonomous systems want millisecond feedback.

Severity: Medium-High. The coping strategy is to move collection into commodity devices and browsers, then add verification and filtering later, but that still leaves an unsolved trust-versus-latency tradeoff that looks worth building for.


3. What People Wish Existed

Workflow-grounded evaluation that model labs cannot ignore

The clearest need was for evaluation systems that speak the language of operations, not just leaderboards. @OfficialLoganK said (37 likes, 8 replies, 2,365 views) companies should spend a quarter of their time on benchmarks and get model labs to care about them, while @MatznerJon asked (26 likes, 9 replies, 4,432 views) whether any new AI technique changes contribution margin or on-time task completion. @DataChaz pointed (22 likes, 1 reply, 1,178 views) to Codos precisely because benchmark wins are not enough for enterprises. Practical urgency is high, and partial answers exist in products like Devin's FrontierCode and Codos-style workflow measurement, but the feed still treats the gap as wide open. Opportunity: direct.

Typed decision layers with logs, confidence, and safe fallbacks

What people want from Jev-style systems is not generic "smarter agents." They want a decision layer that can route, score, approve, or reject with explicit confidence, while every risky or deterministic rule stays inspectable in code. @0xwhrrari framed (25 likes, 8 replies, 524 views, 21 bookmarks) the architecture, @0xClodex listed (8 likes, 1 reply, 216 views, 10 bookmarks) the operating rules, and @ClementDelangue argued (84 likes, 13 replies, 6,554 views, 19 bookmarks) that many production use cases do not need full LLM APIs at all. This is a practical need with immediate ROI pressure rather than an aspirational wish. Opportunity: direct.

Post-training and harness tooling that preserves standard model formats

The Halo and ModularRSI threads point to a more technical but very concrete need: teams want scale-up tooling that does not force them to rewrite models, lose standard checkpoints, or mix genuine harness improvement with benchmark-fitting. @whitecircle positioned (171 likes, 66 replies, 15,372 views, 120 bookmarks) Halo as the bridge between stock TRL and much heavier infrastructure, while @dair_ai highlighted (52 likes, 14 replies, 3,731 views, 37 bookmarks) ModularRSI because reusable harness gains are still hard to isolate. The need is practical, but the buyer is a more technical team with some existing infrastructure maturity. Opportunity: competitive.

Queryable spatial ground truth with tunable trust and latency

Physical-AI posters were effectively asking for a world-state API that is cheap to collect, easy to query, and fast enough for robots. @HuuHoang88 wanted (22 likes, 24 replies, 205 views) commodity-device capture instead of expensive fleets, @KaiBGR wanted (14 likes, 16 replies, 1,206 views) verification without unbearable delay, and @itsbac0201 highlighted (12 likes, 9 replies) browser teleoperation as the current workaround. Vangrid and AXIS partially address the need today, but both still read as early infrastructure rather than finished defaults. Opportunity: competitive.


4. Tools and Methods in Use

Tool Category Sentiment Strengths Limitations
MiMo-V2.6 Pro Omnimodal open model (+) 46 AA index, strong DeepSWE/AutomationBench rows, open RL assets, strong cost frontier 573 GB-class repo footprint, still behind frontier peers on some rows such as ProgramBench, Toolathlon, and Terminal Bench 4.0
MiMo-V2.6 Flash Omnimodal open model (+) 15B active parameters, 1M context, lower cost, close to Pro on several agent tasks Still roughly 178 GB on disk, trails Pro and some closed models on multiple benchmarks
Grok 4.7 Frontier model (+/-) Lower price than GPT-5.6 Sol Max and Fable 5.1 Max, strong DeepSWE and office-work claims, available in Devin Vendor-linked benchmark skepticism, over-scoping in some engineering tasks, weaker Terminal Bench 4.0 row
Jev Decision / classifier model (+) Typed routing, scoring, and verification; confidence scores; cheaper repeated decisions Requires explicit criteria and logging, cannot replace text generation, can add composition complexity
Halo Training framework (+) Hugging Face-native distributed training, async RL, same checkpoint format, 2.3-2.8x over stock TRL claims Requires more GPU and infra maturity than stock prototype tooling
TRL Training framework (+/-) Familiar Hugging Face starting point for small or early runs Throughput and memory ceiling for larger or MoE post-training, often forces a later stack change
Codos Workflow automation platform (+/-) Maps real work with AI interviews, ties rollout to measurable workflow outcomes, keeps company context in the loop Service-heavy deployment model, deep integration burden for each customer
AXIS Robot data engine (+) Browser teleoperation, quality refinement, augmentation, fixed evaluation snapshots, growing manipulation dataset Simulation-to-reality transfer and backend curation are still major burdens
Vangrid Enterprise Spatial API Spatial data API (+/-) Query/ingest/stream live ground-truth observations with commodity-device capture Early-access product, and verification latency can conflict with millisecond robotics use cases

Overall sentiment split by layer. MiMo and parts of Grok were praised when they moved the cost-performance frontier, not simply when they posted a good number; Jev was attractive because it carved out a cheaper decision layer, but serious posts immediately wrapped it in confidence gates, logs, and hard-coded rules; and Halo plus ModularRSI were valued because they reduce rewrites and make harness behavior easier to reason about. The migration pattern was job-class routing: open-weight models or typed-decision systems for repeated, cost-sensitive steps, larger frontier models for harder generation or unresolved cases, and custom benchmarks to decide where that boundary should sit.


5. What People Are Building

Project Who built it What it does Problem it solves Stack Stage Links
MiMo-V2.6 Pro / Flash @XiaomiMiMo Open-source omnimodal reasoning models for coding, agents, 3D, research, and computer use Narrows the open-weight capability/cost gap for long-horizon multimodal work Sparse MoE, scaled RL, multimodal IO, tool use, 1M context Shipped tweet, release
Halo @whitecircle Hugging Face-native distributed framework for pre-training through async RL Avoids rewriting model families just to scale post-training PyTorch, FSDP2, DTensor, DeepEP, FlashAttention, Liger, Ray, vLLM/SGLang Shipped tweet, repo
ModularRSI IQuestLab Modular harness self-improvement framework for long-horizon agents Separates reusable harness gains from benchmark-fitting and one-off prompt hacks Five harness modules, benchmark-disjoint evolution pool, contrastive trajectory analysis, validation gates Alpha tweet, repo, paper
Codos Codos Virtual Chief AI Officer platform that maps work and rebuilds workflows around AI Turns benchmark chatter into measurable capacity, speed, or revenue changes AI interviews, company context graph, embedded engineers, workflow apps Shipped quote, site
OpenClaw OpenClaw Self-hosted assistant that keeps working across chat apps and background jobs Lets agents stay on after the laptop closes and operate through WhatsApp or similar channels Node.js, browser control, shell access, persistent memory, plugins Shipped tweet, repo
phantom-kv @lordx64 Loadable KV-cache graft for reversible refusal removal Changes model behavior without permanent checkpoint edits or model reloads Python, PyTorch, Transformers, per-layer KV tensors Alpha tweet, repo
AXIS data engine Axis Robotics Browser teleoperation pipeline plus manipulation dataset and benchmark Scales robot-data collection without specialized hardware or local simulator installs MuJoCo-WASM, backend validation, smoothing, augmentation, VLA training Beta tweet, site
Vangrid Enterprise Spatial API Vangrid API for spatial queries, ingestion, and live ground-truth streams Supplies physical-AI systems with current real-world state instead of static maps or pure simulation Edge captures, verification pipeline, REST + streaming Beta tweet, API docs

MiMo and Halo were the clearest infrastructure-side launches on the software half of the feed. MiMo pushes open model supply outward with enough capability and pricing pressure to force live comparisons against frontier APIs, while Halo attacks the operational cost of getting from Hugging Face prototypes to serious post-training.

ModularRSI and phantom-kv show a second builder pattern: people are modifying the control substrate around models rather than only shipping another chat surface. One makes harness improvement more modular and benchmark-disjoint; the other makes model behavior switchable without touching base weights.

Codos, OpenClaw, AXIS, and Vangrid all attack bottlenecks outside the model itself: company workflow mapping, persistent execution, manipulation-data collection, and live spatial truth. The repeated trigger across these builds is the same one seen elsewhere in the report: deterministic scaffolding, verified state, and workflow measurement matter more than adding one more undifferentiated model call.


6. New and Notable

Efficient LLM design became a formal course topic

@aminkarbasi announced (63 likes, 4 replies, 5,685 views, 64 bookmarks) Stanford's MS&E 319 course on Efficient Generative Language Models, explicitly centering compute budgets, training objectives, model architecture, and inference algorithms in one syllabus. The topic list mattered because it matched the day's live debates almost exactly: MoE design, KV-cache compression, quantization, speculative decoding, LoRA, RLHF, DPO, and distillation all sat in the same frame as MiMo and Halo.

Independent evaluation moved into the mainstream governance feed

@Yoshua_Bengio welcomed (148 likes, 16 replies, 4,764 views) calls for mandatory pre-deployment testing, independent evaluation, common standards, and an intergovernmental organization for frontier AI. That mattered because evaluation was not only a builder complaint on 2026-09-21; it also appeared as a policy priority at the top of the feed.

Another giant sparse open-weight model was already being pre-sold for local deployment debates

@TeksEdge previewed (29 likes, 3 replies, 1,775 views, 7 bookmarks) "Step 5" as a 600B-parameter, 27B-active, 1M-context open-weight model due on October 15. The notable part was not just the teaser. It was the deployment question underneath it: if Kimi K3-class intelligence can come in a much smaller sparse package, local hardware planning may shift from sheer total parameters toward tiered memory and active-parameter efficiency.


7. Where the Opportunities Are

[+++] Workflow-grounded evaluation and routing control β€” Evidence from sections 1-4 kept converging on the same gap: @MatznerJon asked for contribution-margin and on-time-task-completion proof, @OfficialLoganK pushed for benchmark budgets, @cognition showed how benchmark gains can disappear on specific task shapes, and Jev posts kept separating cheap decisions from expensive generation. This is strong because the demand appears in frustration posts, tool releases, and shipped products like Codos.

[++] Hugging Face-native post-training and harness infrastructure β€” Halo and ModularRSI attack the middle ground where teams outgrow stock TRL but do not want to rebuild everything. The opportunity is moderate because the pain is acute and technical, but the buyer is narrower and more sophisticated than for evaluation tooling.

[++] Physical-world data rails with tunable trust versus latency β€” AXIS, Vangrid, and the surrounding discussion all point to the same missing layer: affordable collection plus enough verification to trust the feed without making it too late for real robots to use. This is moderate because the bottleneck is real and persistent, but integration and go-to-market surfaces are still early.

[+] Hot-swappable model behavior layers β€” phantom-kv shows one version of reversible per-request capability overlays that do not rewrite base weights, and Jev-style control layers show another version of moving behavior outside the main generator. This is emerging rather than mature, but the pattern is real enough to watch.


8. Takeaways

  1. MiMo, not generic open-source churn, was the day's main model event. Xiaomi paired frontier-adjacent agent scores with open RL assets and cost pressure on closed APIs, which is why the conversation immediately moved from announcement hype to deployment math. (source)
  2. Grok 4.7 got attention because of price-adjusted engineering value, not because everyone agreed it was the new undisputed frontier. Cognition's own deployment notes said the model is strong on hard backend tasks but can still trail Grok 4.6 overall when it over-scopes. (source)
  3. The feed increasingly wants agents to generate, decide, and verify in different layers. Jev posts kept repeating the same split: LLMs for creation, typed decision systems for routing/scoring, deterministic code for limits, and humans for uncertainty. (source)
  4. Builder energy is moving below the chat surface into training, harnesses, and workflow instrumentation. Halo, ModularRSI, and Codos all attack infrastructure or operating-model problems rather than just shipping another assistant wrapper. (source)
  5. Physical AI still looks more data-constrained than model-constrained. AXIS and Vangrid examples focused on browser teleoperation, phone-based capture, verification, and live world-state feeds rather than on new robot hardware. (source)