Twitter AI - 2026-10-08¶
1. What People Are Talking About¶
1.1 Evaluation moved from static scoreboards to maintained systems (🡕)¶
Among the strongest evidence-rich items in the feed, evaluation infrastructure was the clearest cluster. At least nine retained items focused on keeping benchmarks fresh, hard to game, and close to real work: agentic web-search evals, terminal-agent RL environments, dynamic robotics benchmarks, microbenchmarks, double-blind evaluation, manual judging for open-ended work, and judge-model disagreement. The recurring message was that a single frozen leaderboard is no longer enough.
@ExaAILabs introduced (61 likes, 5 replies, 1,867 views, 13 bookmarks) ATLAS as a benchmark for how agents search now, not how older evals assume they search. The linked ATLAS write-up says 74% of tasks require finding 10 or more entities, full discovery plus enrichment needs a median of 18 different domains, and degraded search quality cuts ATLAS scores far more sharply than BrowseComp, WideSearch, or DeepSearchQA. That makes the benchmark's thesis concrete: if search quality gets worse, the score should fall hard.

@Weyaxi released (38 likes, 3 replies, 1,760 views, 26 bookmarks) TermGrade, an open set of RL environments for terminal agents. The tweet and thread are unusually specific: 1,000 executable environments, 36,000 trajectories, only 1.5% of 66,000 generated candidate tasks surviving validation, and training concentrated on tasks where the base model succeeds about half the time. That is not a generic "we made a benchmark" post; it is a statement about task quality control and reward-shaping discipline.
@SnorkelAI announced (20 likes, 628 views) that Open Benchmarks Grants is expanding to a $30M commitment, adding a benchmark red team and a research fellowship. The linked announcement is explicit about the failure modes it wants to attack: contamination, broken verifiers, reward hacking, and benchmark staleness across safety, cybersecurity, physical AI, long-running agents, and open-ended outputs.
@Promzy__M argued (24 likes, 32 replies, 350 views) that frozen robotics benchmarks stop telling the truth once models start memorizing them. The attached Open Axis images matter because they show a live benchmark surface rather than a slogan, while the linked AXIS site says the system currently spans 207 tasks and 50K+ trajectories with a held-out protocol and browser-based teleoperation feeding the data engine.

@arpit_bhayani argued (30 likes, 5 replies, 1,637 views, 12 bookmarks) that microbenchmarks belong beside end-to-end evals because they isolate why a system changed: prompt drafting time, vector-search latency, MCP/API waits, embedding throughput, cache hit rates, and memory per concurrent request. @MTSlive summarized (10 likes, 1 reply, 2,748 views, 8 bookmarks) Andrew Trask's case for double-blind frontier-model evaluation via confidential compute, so labs can be evaluated on tests they never see. @kellyhongsn highlighted (14 likes, 1 reply, 598 views) Epoch's choice to manually judge realistic open-ended tasks because automatic grading misses capability gaps often described as a lack of taste. @ArtificialAnlys showed (14 likes, 1 reply, 1,114 views) that six hallucination checkers can uphold very different numbers of errors on the same deliverables, which turns evaluator choice into a real measurement variable.
Discussion insight: Replies and companion posts kept returning to the same objection: stale tasks, overfit verifiers, and opaque judges produce scores that look precise but are not trustworthy. The practical responses on display were red-teaming, private test sets, manual grading for open-ended work, and more granular component-level measurement.
Comparison to prior day: Compared with 2026-10-07, the benchmark conversation became even more operational. The feed moved from broad benchmark funding and dynamic-eval rhetoric toward mechanics such as degraded-search ablations, task survival rates, double-blind testing, and judge-model disagreement.
1.2 Cheap decision layers and routing logic became a first-class product surface (🡕)¶
A second cluster focused on the layer before the frontier-model call. At least five retained items argued that many production decisions should be typed, cheap, and probability-bearing rather than pushed through a general-purpose reasoning model every time.
@ericwilliamrea said (75 likes, 5 replies, 11,418 views, 20 bookmarks) Podium has been alpha testing OpenAI's Decisions API to route agents into the right workflow faster and more cheaply. The quoted launch tweet says the API can choose the right model, tool, or action up to 10x faster than GPT-6 Luna through the Responses API, which is exactly the sort of structured pre-routing layer several other posts were circling.
@RubricLabs shared (18 likes, 5 replies, 1,058 views, 9 bookmarks) practical Jev use cases, and the linked Jev post explains the model as a few-hundred-millisecond decision engine for choice, score, and boolean questions. The examples are operational rather than theoretical: notification prioritization, semantic search over a fixed sentence set, spelling checks, settings navigation, and iterative rewrite loops.
The most concrete practitioner thread came from @jurbed, who documented (4 likes, 1 reply, 222 views, 4 bookmarks) five live Jev deployments. The strongest example was a Hermes firewall that asks whether untrusted text contains commands unrelated to the content and aimed at an AI; the thread says it caught 89% of prompt-injection attacks at a 3.5% false-positive rate on a 718-item benchmark, at about $0.035 per thousand scans. Even at low engagement, this was one of the day's most detailed pieces of first-hand deployment evidence.
@0xwhrrari argued (47 likes, 9 replies, 746 views, 33 bookmarks) that GPT-6 Astra should be reserved for hard tasks because using it everywhere gets expensive fast. The linked x.com article was unavailable to fetch, but the tweet itself, its save-heavy bookmark count, and its fit with the Decisions/Jev discussion made the cost-routing argument clear enough: model choice is now a workflow design problem, not just a benchmark winner problem.
Discussion insight: The main pushback was not against smaller decision models themselves; it was against invisible failure. Replies wanted confidence thresholds, low-confidence escalation, memory-boundary discipline, and proof that a cheap routing layer does not quietly distort the workflow it is supposed to optimize.
Comparison to prior day: The prior day already treated harness design as important. On 2026-10-08, that theme became more concrete through live deployments, cost gating, and a production user describing a structured routing layer in front of full agents.
1.3 Builders kept working on the plumbing around agents: harnesses, local serving, file normalization, and on-device memory (🡕)¶
Another clear theme was that builders are spending energy on the boring parts of agent systems: lighter harnesses, shared local inference, safer file ingestion, and persistent local memory. At least six retained items fit this cluster, and most of them were about cost, control, or operability rather than raw model capability.
@businessbarista outlined (31 likes, 12 replies, 3,705 views, 74 bookmarks) a 5-day enterprise agent sprint that starts with process mapping, alias resolution, and mocked integrations before any production connection. The replies sharpened the operational flavor of the post: one reader said alias mapping is where preventable mistakes show up, another said the process should be fixed before automating it, and another suggested contract tests before swapping emulator paths for real systems.
@upperwal open-sourced (6 likes, 3 replies, 183 views, 6 bookmarks) Loop, a Rust agent harness billed as 44x lighter than opencode at idle and 5x lighter on CPU. The reply thread is where the useful evidence lives: 13/13 tasks passed in the posted suite, median task time dropped to 29s from OpenCode's 72s, and token use fell to 803k from OpenCode's 1.82M.
@basecmpt open-sourced (6 likes, 4 replies, 204 views) Superfluid, a local-AI-native LLM server for multi-agent workloads. The tweet is highly architectural: one scheduler shared across lanes, shared prefix caching, durable session logs that survive worker crashes, runtime-agnostic backends, and the ability to distribute sessions across multiple machines on a LAN.
@Krivoblotsky released (11 likes, 2 replies, 50,480 views) public benchmarks for MacPaw's on-device AI stack. The linked Elix page says inference runs on the Mac with no network path, while Mnemos is a local memory layer that stores facts, files, and history in a citation-bearing knowledge graph instead of a pile of text chunks.

@Muzammil_OX highlighted (8 likes, 8 replies, 145 views) Microsoft's MarkItDown, which converts PDFs, Office files, images, audio, HTML, ZIPs, and more into Markdown for LLM pipelines. The repository page also warns that it performs I/O with the current process's privileges, which is a useful reminder that even the ingestion layer needs security thinking. @NaceAI introduced (32 likes, 9 replies, 3,521 views, 22 bookmarks) NDI 1.0, a small document model trained on 15M+ financial files and marketed as a cheaper parsing option than GPT-6 Astra, though that cost/performance story is still vendor-reported inside the dataset.
Discussion insight: The strongest plumbing posts won attention when they made one of three promises concrete: fewer tokens, lighter local resource use, or less fragile input handling. The common direction was away from monolithic "one big agent stack" thinking and toward layered systems with explicit schedulers, caches, file converters, and narrow specialist models.
Comparison to prior day: This continues 2026-10-07's workflow-layer pattern, but with more concrete artifacts. The prior day emphasized decision layers and eval harnesses; today added lighter coding harnesses, local serving, document conversion, and public on-device runtime benchmarks.
1.4 Local-language and public-sector AI showed up as deployment work, with obvious integration gaps (🡒)¶
A smaller but concrete theme was regional-language AI moving from broad aspiration into deployment and ecosystem-building. The strongest examples were not about new base models; they were about what local-language systems must expose for outside builders and public-service teams to actually use them.
@bosuntijani called for builders (126 likes, 4 replies, 4,782 views, 80 bookmarks) to build on N-ATLAS. The attached poster made the deadline and official site visible, and the fetched N-ATLAS page says it is Nigeria's first open-source multilingual LLM, fine-tuned from Llama-3 8B with Yoruba, Hausa, Igbo, Nigerian-accented English, and ASR support. The immediate replies were practical: where is the data resident, and is there a REST API?

@PriyankKharge said (51 likes, 5 replies, 1,503 views) the BHASHINI Rajyam workshop in Bengaluru brought government departments, institutions, technical teams, and startups together around language AI for public-service delivery. The tweet is the main evidence here: speech-to-text, text-to-speech, translation, transliteration, OCR, and language detection across 23+ Indian languages and 65+ dialects, plus explicit work on Karnataka-specific datasets and glossaries.
Discussion insight: The friction was not ideological. It was operational: API surfaces, data residency, glossary quality, and whether teams on different stacks can actually plug into the platform.
Comparison to prior day: This was more deployment-facing than the prior day's feed, which was dominated by benchmark design, routing layers, and workflow control planes.
2. What Frustrates People¶
Benchmark trust now breaks at the evaluator layer¶
Severity: High. The feed repeatedly showed that evaluation failure is no longer just about missing one more benchmark. @ExaAILabs showed (61 likes, 5 replies, 1,867 views, 13 bookmarks) that degraded search should sharply degrade a search benchmark score, because otherwise the benchmark is not actually measuring search. @SnorkelAI said (20 likes, 628 views) benchmark builders now need active defense against contamination, reward hacking, and broken verifiers. @ArtificialAnlys showed (14 likes, 1 reply, 1,114 views) that different hallucination judges uphold very different error counts on the same deliverables, while @MTSlive summarized (10 likes, 1 reply, 2,748 views, 8 bookmarks) Andrew Trask's argument for double-blind evaluation because labs otherwise hill-climb on tests they can inspect. People are coping with benchmark red-teams, manual grading, microbenchmarks, and private test sets. Worth building: High.
Human review is still the bottleneck in AI-assisted security and QA¶
Severity: High. The starkest hard number came from Anthropic's OSS Scanner launch: the company says it found 29,000+ candidate vulnerabilities but manually reviewed only about 6,000. @AnthropicAI launched (166 likes, 13 replies, 17,075 views, 25 bookmarks) the service precisely because maintainers increasingly want faster reports than humans can validate. That same bottleneck showed up in a different form in @ArtificialAnlys, where judge choice changes the apparent hallucination count. Teams are coping by accepting raw model-generated findings, narrowing scope, and keeping humans in the loop for high-severity or ambiguous calls. Worth building: High.
Frontier models are too expensive or too heavy to be the default for every step¶
Severity: High. @0xwhrrari argued (47 likes, 9 replies, 746 views, 33 bookmarks) that GPT-6 Astra should be saved for hard tasks, and @ericwilliamrea said (75 likes, 5 replies, 11,418 views, 20 bookmarks) a structured Decisions API can route work faster and more cheaply than a full LLM pass. On the harness side, @upperwal posted (6 likes, 3 replies, 183 views, 6 bookmarks) benchmark claims that Loop used less than half the tokens of OpenCode in the posted suite. The coping strategies on display were routing, smaller decision models, lighter harnesses, and more deliberate demo-first prototyping instead of immediate production integration. Worth building: High.
Local-language and public-sector AI still lacks boring but essential integration details¶
Severity: Medium-High. @bosuntijani promoted (126 likes, 4 replies, 4,782 views, 80 bookmarks) N-ATLAS, but the most substantive replies immediately asked about data residency and the absence of a REST API. @PriyankKharge described (51 likes, 5 replies, 1,503 views) BHASHINI work that still needs better local datasets and glossaries for Kannada and other languages. Even outside language infrastructure, @Muzammil_OX highlighted (8 likes, 8 replies, 145 views) MarkItDown because getting messy files into LLM-ready form remains a real workflow problem. Worth building: High.
3. What People Wish Existed¶
Trustworthy dynamic evals with private tests and judge QA¶
People increasingly want evaluations that move with the frontier, resist memorization, and explain failure modes instead of only producing a rank. @SnorkelAI and @ExaAILabs both pushed in that direction, while @MTSlive and @maishimadamd added private-test and dynamic-rubric nuance. This is a practical need with multiple visible failure cases today. Opportunity: Direct.
Typed routing and workflow firewalls that can sit in front of agents¶
The feed repeatedly asked for cheaper, structured decisions before the expensive model call. @ericwilliamrea showed production routing demand, @RubricLabs documented a model built for typed decisions, and @jurbed showed how that layer can also filter prompt-injection risk. Some solutions exist already, but the market still looks early enough that teams are piecing together custom stacks. Opportunity: Direct.
Local-first runtime and ingestion layers for multi-agent work¶
A clear wish hiding inside multiple launches is for infrastructure that assumes several agents are running at once on constrained machines, with private local inference and sane file handling. @upperwal, @basecmpt, @Krivoblotsky, and @Muzammil_OX each addressed part of that stack. This need is concrete, but several builders are already competing around it. Opportunity: Competitive.
Public-sector language AI with real APIs, residency guarantees, and glossary tooling¶
The regional-language posts did not read like abstract policy. They read like teams asking for implementation surfaces: APIs, deployment clarity, local datasets, and domain glossaries. @bosuntijani and @PriyankKharge showed that interest directly. The need looks practical and urgent for specific regions, but the buyer and procurement environment will be narrower than general developer tooling. Opportunity: Direct.
4. Tools and Methods in Use¶
| Tool | Category | Sentiment | Strengths | Limitations |
|---|---|---|---|---|
| ATLAS | Search benchmark | (+) | Measures search quality directly, uses wide/deep multi-domain tasks, penalizes degraded retrieval | Still one benchmark surface and requires a complex harness |
| TermGrade | RL environment / benchmark | (+) | 1,000 executable environments, validated tasks, open trajectories, grading tied to execution | Early release with builder-reported gains |
| Open Benchmarks Grants | Evaluation program | (+) | Funds benchmark creation, adds red-teaming, supports fellowship research | It is supporting infrastructure, not a turnkey harness |
| Open Axis Benchmark / AXIS | Robotics benchmark + data engine | (+) | Moving task library, browser teleoperation, held-out protocol, robustness gains under perturbation | Specialized to tabletop robotics workflows |
| OpenAI Decisions API | Routing / decision model | (+/-) | Structured model/tool/action choice, faster and cheaper than a full LLM pass for routing | Users still question workflow integrity and memory boundaries |
| Jev | Decision model | (+) | Typed outputs, calibrated probabilities, few-hundred-millisecond decisions, real deployment evidence | Requires predefined answer spaces and threshold tuning |
| Loop | Agent harness | (+/-) | Low idle overhead, lower CPU time, lower token use in posted benchmarks | Developer preview and benchmark evidence is builder-reported |
| Superfluid | Local LLM server | (+/-) | Shared scheduler, shared prefix cache, crash-tolerant sessions, multi-machine support | Evidence today is mostly self-reported architecture claims |
| Elix + Mnemos | On-device runtime + memory | (+) | No network path for inference, portable local memory, citation-bearing knowledge graph | Runtime remains closed source and both products are still in development |
| MarkItDown | Document ingestion | (+) | Broad file support, Markdown output aligned with LLM pipelines, CLI and library modes | Repository explicitly warns about input trust and process-level I/O access |
| N-ATLAS | Local-language LLM | (+/-) | Open-source positioning, Nigerian language support, ASR for voice workflows | API and data-residency questions surfaced immediately |
| BHASHINI | Public-sector language platform | (+) | Broad language and dialect coverage, speech and translation features for service delivery | Still depends on glossary and dataset work for local deployment quality |
| NDI 1.0 | Document model | (+/-) | Smaller-model cost story for parsing/classifying financial documents | Benchmark framing is vendor-reported in this dataset |
The highest-confidence satisfaction signals went to tools that made evaluation or routing more inspectable: ATLAS, Jev, Snorkel's red-team framing, and AXIS. Sentiment turned mixed when the evidence came mainly from a builder's own benchmark thread, as with Loop, Superfluid, and NDI 1.0.
The clearest migration pattern was away from "send everything to one frontier model" and toward layered systems: typed decisions first, heavier reasoning second; local serving and shared caches instead of per-agent inference silos; specialist document models or converters instead of forcing every file through a general chat model. Evaluation followed the same pattern, shifting from frozen public scoreboards toward dynamic tasks, manual judging, confidential compute, and benchmark maintenance.
5. What People Are Building¶
| Project | Who built it | What it does | Problem it solves | Stack | Stage | Links |
|---|---|---|---|---|---|---|
| OSS Scanner | Anthropic | Opt-in vulnerability scanner for open-source repositories that sends model-generated reports with reproducers and proposed patches | Human security teams cannot review growing vulnerability volume fast enough | Frontier Claude/Mythos models, isolated VMs, offline scanning, GitHub enrollment repo | Beta | launch post, GitHub |
| ATLAS | @ExaAILabs | Benchmark for how agents actually search and enrich results across many domains | Older search benchmarks are outdated, contaminated, or solvable without real search | Search harness, Discovery/Row/Item F1 metrics, multi-domain tasks | Shipped | tweet, benchmark |
| TermGrade | @Weyaxi | Open RL environments plus trajectories and grading recipe for terminal agents | Terminal agents need executable environments, not just prompt-only tasks | 1,000 executable environments, 36,000 trajectories, execution-based grading | Beta | tweet |
| Loop | @upperwal | Rust coding-agent harness designed to be lighter than opencode | Existing harnesses waste tokens and local CPU just sitting idle | Rust, CLI harness, posted GLM 5.3 Flash benchmark suite | Alpha | tweet |
| Superfluid | @basecmpt | Local-AI-native LLM server for multi-agent workloads on constrained machines | General LLM servers are not tuned for many local agents sharing one device | Shared scheduler, shared prefix cache, runtime-agnostic workers, multi-machine placement | Alpha | tweet |
| Elix + Mnemos | MacPaw | On-device runtime plus memory layer for Mac assistants | Need private local inference and lasting memory without a cloud dependency | Swift, MLX-based inference, composable pipelines, knowledge-graph memory | Alpha | Elix, Mnemos, tweet |
| NDI 1.0 | @NaceAI | Small model for parsing and classifying financial documents | Repeated document-processing tasks are too expensive for frontier models | Small specialized model, 15M+ financial-file training corpus, benchmarked integrations | Beta | tweet |
| N-ATLAS | NCAIR / NITDA | Multilingual Nigerian LLM platform being pushed into downstream app development | Need locally relevant language AI and voice tooling for Nigerian users | Llama-3 8B fine-tune, multilingual support, ASR | Shipped | tweet, site |
OSS Scanner and TermGrade represent one build pattern: turn previously internal evaluation or security workflows into reusable public infrastructure, even if the human bottleneck has not disappeared. ATLAS and AXIS point in the same direction from a benchmark angle, where the product is not only a score but a maintained task surface.
Loop, Superfluid, and Elix/Mnemos reflect a second pattern: builders are optimizing the runtime around the model rather than only the model itself. The recurring pain points are idle overhead, shared scheduling, privacy, persistence, and the cost of running multiple agents on one machine.
NDI 1.0 and N-ATLAS show a third pattern: smaller or more local models targeted at a specific deployment surface. One is vertical and economic, aimed at repeated financial-document work; the other is geographic and linguistic, aimed at local relevance and voice-heavy language access.
6. New and Notable¶
Anthropic turned frontier vulnerability finding into an opt-in OSS service¶
@AnthropicAI launched (166 likes, 13 replies, 17,075 views, 25 bookmarks) OSS Scanner as a no-cost service for opted-in open-source projects. The linked launch post says Anthropic has found 29,000+ candidate vulnerabilities but manually reviewed only about 6,000, and that early OSS Scanner reports include reproducers and candidate patches generated without human review. That makes the launch notable both as a product and as a public admission that validation capacity is the bottleneck.
Benchmark builders kept upgrading the benchmark itself, not just the scorecard¶
@ExaAILabs introduced (61 likes, 5 replies, 1,867 views, 13 bookmarks) ATLAS to punish weak search, while @SnorkelAI expanded (20 likes, 628 views) its open-benchmark effort with a red team and fellowship. @Promzy__M used (24 likes, 32 replies, 350 views) Open Axis to argue that robot benchmarks need moving task sets, and @ArtificialAnlys showed (14 likes, 1 reply, 1,114 views) that the hallucination judge itself can change the outcome.

MacPaw made its on-device stack legible through public benchmark pages¶
@Krivoblotsky surfaced (11 likes, 2 replies, 50,480 views) benchmark pages for Elix and Mnemos, turning an on-device runtime and memory layer into something observers can inspect instead of only imagine. The Elix page is especially notable because it claims no network path from inference and publishes comparative throughput charts against other Apple-silicon runtimes.
N-ATLAS and BHASHINI made local-language deployment feel operational¶
@bosuntijani promoted (126 likes, 4 replies, 4,782 views, 80 bookmarks) a live challenge around N-ATLAS, while @PriyankKharge described (51 likes, 5 replies, 1,503 views) state-focused BHASHINI workshops around service delivery, local glossaries, and dataset work. These were notable because they framed language AI as implementation and procurement work, not just model nationalism.
7. Where the Opportunities Are¶
[+++] Dynamic evaluation operations and audit tooling - Evidence spans ATLAS, Snorkel's benchmark red-team push, Open Axis, TermGrade, Epoch's manual evaluation note, Andrew Trask's double-blind evaluation argument, and the hallucination-judge comparison. This is strong because multiple independent posts pointed to different parts of the same trust problem.
[+++] Typed routing, workflow guardrails, and AI-firewall layers - Decisions API, Jev's public use cases, jurbed's five deployments, and Astra cost gating all point to a shared need: fast structured decisions in front of the expensive model call, plus confidence-aware escalation and prompt-injection filtering.
[++] Local-first agent runtime and document plumbing - Loop, Superfluid, Elix/Mnemos, MarkItDown, and NDI 1.0 all target the same broad surface: make agents cheaper to run, easier to feed, and less dependent on remote infrastructure. The signal is real, but several claims are still early or builder-reported.
[++] Public-sector and regional-language AI enablement - N-ATLAS and BHASHINI both surfaced practical needs around APIs, residency, glossaries, and local datasets. The demand looks concrete, especially for governments and local ecosystems, but the market is narrower and more deployment-specific.
[+] Security triage and patch-verification workflows for model-generated findings - OSS Scanner exposed a large gap between candidate findings and human review capacity. The opportunity is emerging because the evidence is strong, but the operating model is still being worked out in public.
8. Takeaways¶
- The strongest evidence-rich part of AI Twitter was about evaluation mechanics, not model hype. ATLAS, Open Axis, Snorkel's red-team expansion, microbenchmarks, and manual judging all focused on how to keep measurement honest as systems improve. (ATLAS, Snorkel, arpit_bhayani tweet)
- Typed decision layers are moving into real production roles. Decisions API, Jev's published examples, and jurbed's deployment thread all treat routing, filtering, and classification as their own product layer rather than a side effect of a chat model. (ericwilliamrea tweet, Jev post, jurbed tweet)
- Builders are competing on the plumbing around agents: harness weight, local serving, document conversion, and memory. Loop, Superfluid, MarkItDown, and Elix/Mnemos all target operating cost and control more than raw frontier capability. (upperwal tweet, basecmpt tweet, MarkItDown, Elix)
- Local-language AI showed real deployment pull, but integration details are still the friction. N-ATLAS and BHASHINI both surfaced demand for local relevance, while replies and workshop goals highlighted API, residency, dataset, and glossary gaps. (N-ATLAS site, bosuntijani tweet, PriyankKharge tweet)
- Human review remains the bottleneck even when AI produces more candidate findings or more judge output. Anthropic's 29,000+ candidate vulnerabilities versus about 6,000 manually reviewed, plus judge-model disagreement in hallucination checking, both point to review throughput as a core constraint. (OSS Scanner launch, Artificial Analysis tweet)