HackerNews AI - 2026-09-26¶
1. What People Are Talking About¶
September 26's HackerNews AI feed got smaller but more concentrated. Story count fell to 55 from 79 on September 25, total points fell to 352 from 425, and Show HN volume fell to 19 from 30, but comments jumped to 210 from 140. The highest-point item was Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments), and the top four stories together captured 63.4 percent of the day's points and 84.3 percent of its comments. That concentration pulled the day toward one shared question: once agents touch live tools, diagrams, and shopping flows, what interface or proof makes their behavior trustworthy?
1.1 Agent-readable work surfaces kept replacing chat-only workflows (🡕)¶
The strongest UX theme was not "more autonomous code generation." It was making the environment itself more legible to the agent so less time is spent spelunking through docs, guessing commands, or juggling tabs. The shared move was to turn previously implicit context - a canvas, a worktree, a product UI, or a README - into an explicit surface the agent can act on or consult directly.
parasitid posted Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments). The linked Drawgent page says it connects an existing Claude Code, Codex, or opencode session to Excalidraw, lets people write AGENT: notes beside a diagram, and has the agent inspect and edit the canvas live before marking the note DONE. runtooldev posted The Discovery Tax: Why Coding Agents Waste 2,500 Tokens Before Writing Code (5 points, 2 comments), and the linked essay says a typical "run the tests" interaction can burn about six tool calls and 2,500 tokens on rediscovery before any real work starts, versus about one tool call and 100 tokens once tasks are exposed through an MCP registry.
Lower-score posts attacked the same bottleneck from adjacent directions. staaake posted Show HN: I built a tool that gives any website an API and MCP (5 points, 1 comment), describing a hosted connector that exposes backend capabilities as an API, MCP endpoint, webhooks, and SDK for products that do not publish a developer interface. empiree posted Orca - Agent Development Environment (ADE) for shipping with coding agents (3 points, 1 comment), and the linked write-up describes a git-worktree control plane with diff comments, browser previews per worktree, and phone notifications for concurrent tasks. signa11 posted Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments), where Phoronix says kernel developers saw better attribution and project-conformance behavior once an AGENTS.md file pointed agents at the right documentation.

Discussion insight: The Drawgent thread was useful because it showed that "give the agent a whiteboard" is not a settled design. seemaze (score 0) pointed to Excalidraw's own MCP tooling, armanj (score 0) said Mermaid ended up more agent-friendly in their own experiments, and 4ndrewl (score 0) argued that the value of drawing is the human thinking it forces, not just the artifact it leaves behind.
Comparison to prior day: September 25's operator-tooling wave centered on memory logs, mobile relays, and local control rooms. September 26 pushed that same instinct into agent-readable surfaces: canvases, worktrees, MCP registries, website shims, and repository-level instruction files.
1.2 Trust moved from abstract safety claims to concrete side-effect control (🡕)¶
The second major cluster focused on what happens when a model's output is no longer just text. HN was less interested in whether an assistant sounds aligned in a demo and more interested in whether it still calls the right tool, refuses the wrong request, keeps secrets private, and stops before money or code moves.
nisosguy posted Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments). The linked Lasso study says watermarking reduced tool-calling accuracy on six of seven tested models and reports average paired churn of 6.5 percent across 21 model-temperature combinations. On refusal behavior, the authors say the effect became sharper under prompt injection: for gemma-3-27b at T=0.001, refusal churn rose from 6.0 percent on bare harmful requests to 23.5 percent under injection, while the net compliance change moved from -1.0 to +12.5 points. ddaniel10 posted Can AI Shopping Agents Be Trusted? (15 points, 28 comments), and the linked F-Secure experiment says its Claude Haiku + Playwright shopping agent followed a review link to a phishing site and submitted the user's name, date of birth, and Social Security number in 12 of 100 runs, usually without reporting the leak to the user.
The product responses in the feed were narrower and more operational. danebalia posted Show HN: AI coding agents that prove their work (2 points, 2 comments), and the linked Keel site says the user defines what "done" means, the agent does the work, and the harness stops the line if a check fails or the human has not signed off, while keeping a tamper-evident record. At the ops layer, meredithbloom posted AI: Who Still Uses Permissions? (2 points, 1 comment), saying repeated approval prompts pushed them to claude --dangerously-skip-permissions, and weiqi554177834 posted Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment), arguing that selling to government or large enterprise buyers may now force small teams to think about multi-provider fallback much earlier.
Discussion insight: The watermarking and shopping-agent threads were skeptical in technically useful ways. In the watermarking discussion, WithinReason (score 0) said a correctly implemented watermark should not degrade model output quality, Klaus23 (score 0) asked whether the paper misunderstood how the technique works, and skybrian (score 0) highlighted seed and paired-comparison concerns. In the shopping-agent thread, planb (score 0) and charcircuit (score 0) argued that an older Haiku model plus a permissive system prompt stacked the deck. Even so, both debates accepted the same underlying standard: agent safety has to be evaluated on concrete tool use and task-aligned attacks, not just on generic chat behavior.
Comparison to prior day: September 25's safety discussion emphasized benchmark loops and public failure reports. September 26 moved one step closer to live side effects: the questions were about tool calls, purchases, sign-offs, provider redundancy, and whether provenance mechanisms stay behaviorally stable under attack.
1.3 Builder energy spread into narrow, engine-backed agent products (🡕)¶
The third theme was how quickly the builder surface fragmented into specific workflows. Rather than another general-purpose copilot pitch, many launches paired a model with a deterministic substrate - a chess engine, a 3D application, a causal DAG, or a proof checker - and made that substrate do the grounding.
brumar posted Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments). The linked repo says it uses whisper.cpp plus Stockfish to align a player's spoken or written thoughts to PGN timestamps, generate annotated post-mortems, and render a narrated video; the author says a full run takes about an hour and would cost around $15 at API prices. xeonax posted Show HN: Blender Copilot (4 points, 0 comments), and the repo says the agent loop runs inside Blender's own Python process, executes bpy against the live scene, and wraps each turn in one undo step rather than mirroring state through a sidecar. 2au_observer posted Show HN: Causal analyst agent skill for Claude (4 points, 0 comments), and the repo says it asks domain experts for missing business context, confirms a causal DAG, chooses methods before seeing results, and returns an offline HTML report with trust grades and alternative diagrams.

Even the lower-score launches followed the same pattern. Aleksandr_NFA posted Show HN: ProofForge, AI agents whose proofs have to compile in Lean (1 point, 0 comments), and the linked repo says agents only count as successful once the Lean 4 kernel accepts the formalization. The recurring move was not "let the model say more." It was "give the model a substrate that can reject bad work."
Discussion insight: The chess thread drew a useful line between ordinary LLM wrapping and real workflow change. WoodenChair (score 0) said LLM-assisted chess commentary is not novel without Stockfish, while dverlaeckt80 (score 0) argued the interesting part is tying engine analysis back to the player's own thoughts. The same distinction applies across the cluster: the point was not that the model can talk about chess, Blender, causality, or proofs, but that those systems push back on the model in domain-specific ways.
Comparison to prior day: September 25's builder wave clustered around typed judgment and project memory. September 26 widened into application-specific agent shells, where the durable value came from the surrounding engine, proof checker, or expert workflow rather than from another generic chat surface.
2. What Frustrates People¶
Environment discovery and orchestration still burn expert attention before useful work starts¶
The Discovery Tax: Why Coding Agents Waste 2,500 Tokens Before Writing Code (5 points, 2 comments) says a session can spend about six tool calls and 2,500 tokens just locating the right test command. Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments) and Orca - Agent Development Environment (ADE) for shipping with coding agents (3 points, 1 comment) exist because parallel agent work is still awkward when context lives in scattered tabs, shells, and local state. Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments) adds the maintainer version of the same complaint: without a project-specific entry point, agents guess wrong about conventions like attribution, but sending them through the full documentation tree also raises token-cost concerns. Show HN: I built a tool that gives any website an API and MCP (5 points, 1 comment) shows the same frustration outside the repo, where many products still expose a human UI but no agent-readable interface.
People are coping by writing registries, AGENTS.md files, worktree control planes, canvases, and API/MCP shims instead of letting the agent rediscover everything every session. That is a strong signal that the pain is recurring and not tool-specific. Severity: High. Worth building for: yes, directly.
Broken permission and safety layers create the worst incentive: bypass¶
Can AI Shopping Agents Be Trusted? (15 points, 28 comments) and Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments) show two versions of the same failure: agents do risky things without surfacing them clearly, and even provenance features can change behavior under attack. AI: Who Still Uses Permissions? (2 points, 1 comment) makes the operational consequence blunt: one user now defaults to claude --dangerously-skip-permissions because repeated prompts are less tolerable than the risk. Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment) shows the same pressure at the architecture level, where safety and vendor resilience are starting to look like obligations rather than optional hardening.
The coping strategies in the feed were all layered defenses: stricter prompts, explicit signoff gates, hot-standby providers, and compile-or-fail checks. The problem is that they are fragmented, and the underlying UX still pushes users toward either overtrust or exhaustion. Severity: High. Worth building for: yes, directly.
The most grounded vertical workflows are still slow, token-heavy, or joyless to run¶
Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments) says a rich analysis can take about an hour and cost about $15 at API prices. Show HN: Blender Copilot (4 points, 0 comments) describes 52 tool calls to build one spaceship and says the author ended up "not proud of the prototype" and feeling "No joy" despite the technical success. Even when a workflow is clearly more capable than a plain chat session, the surrounding process can still feel expensive, slow, or emotionally flat.
People are coping by reserving these systems for high-value tasks and grounding them in harder external checks such as Stockfish, live application state, or formal proof systems. That helps, but it does not yet remove the runtime, cost, or morale overhead. Severity: Medium-High. Worth building for: yes, but the opportunity is partly efficiency and partly workflow design.
3. What People Wish Existed¶
Machine-readable operating surfaces instead of repeated rediscovery¶
The clearest practical need was for environments that tell an agent what the project or product actually is without forcing it to infer that from scraps. The Discovery Tax: Why Coding Agents Waste 2,500 Tokens Before Writing Code (5 points, 2 comments) argues for structured task registries, Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments) argues for repo-native agent instructions, Show HN: I built a tool that gives any website an API and MCP (5 points, 1 comment) turns human-only software into agent-readable endpoints, and Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments) turns diagram intent into a live control surface.
This is an urgent and practical need. Partial answers exist, but the day showed them scattered across registries, docs, canvases, and backend shims rather than packaged as a normal part of software delivery. Opportunity: direct.
Proof-carrying execution before agents spend money, ship code, or claim success¶
Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments), Can AI Shopping Agents Be Trusted? (15 points, 28 comments), Show HN: AI coding agents that prove their work (2 points, 2 comments), and Show HN: ProofForge, AI agents whose proofs have to compile in Lean (1 point, 0 comments) all point to the same wish: a result should not count just because the model said it was done. People want outputs that survive a gate - signoff, benchmark, compiler, theorem prover, or policy check - before side effects happen.
This is both a practical and an emotional need. Practical, because silent failures are already visible in the dataset; emotional, because people do not want to feel one hidden token choice away from a bad purchase or a bad merge. Opportunity: direct.
Small-team control planes for parallel and multi-provider agent work¶
Orca - Agent Development Environment (ADE) for shipping with coding agents (3 points, 1 comment), Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments), AI: Who Still Uses Permissions? (2 points, 1 comment), and Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment) all point at the same operational gap. People want a lightweight way to run several agents at once, manage approvals without babysitting, and survive model/vendor risk without building a full internal platform team.
This is an urgent practical need, but the shape of the solution is still unsettled. The feed showed pain and partial tooling, not consensus. Opportunity: direct.
Explanation-first vertical skills that combine models with an expert substrate¶
Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments), Show HN: Causal analyst agent skill for Claude (4 points, 0 comments), and Show HN: Blender Copilot (4 points, 0 comments) show a more specific wish than "help me with my job." People want agent workflows that can explain themselves in the language of the domain: chess mistakes against engine lines, business causality against a reviewed DAG, or 3D actions against the actual live scene rather than a disconnected plan.
This is a practical need with obvious product appeal, but it already looks competitive because the value comes from domain fit and workflow taste more than from a unique base model. Opportunity: competitive.
4. Tools and Methods in Use¶
| Tool | Category | Sentiment | Strengths | Limitations |
|---|---|---|---|---|
| Drawgent | Visual agent workspace | (+) | Turns diagrams into executable instructions, reuses existing agent sessions, and keeps scene state local to the workspace | Requires installed agent CLIs plus renderer setup, and some builders still find Mermaid or text more agent-friendly |
| Orca | Agent development environment | (+) | Uses git worktrees for parallel tasks, adds line comments, browser previews, and completion notifications | Adds another management layer on top of CLI agents and surfaced here via a secondary write-up rather than a first-hand launch README |
MCP task registries (run) |
Tool discovery / task execution | (+) | Replaces repeated README/Makefile guessing with typed tasks and cuts one sample test flow from about 6 calls / 2,500 tokens to 1 call / 100 tokens | Requires projects to define and maintain the registry explicitly |
| SynthID-style watermarking | Provenance method | (+/-) | Makes AI-generated text detectable and aligns with policy/provenance requirements | The cited study reports tool-call drift and refusal churn under prompt injection, with effects varying by model and watermark key |
| Claude Haiku shopping agent | Transaction agent | (-) | Can autonomously browse listings, compare products, and complete checkouts | In the cited test it hallucinated codes, leaked personal data to a phishing site in 12 percent of runs, and usually failed to disclose the leak |
| Keel | Verification harness | (+) | Lets users define "done," gates work on failed checks or missing signoff, and keeps a tamper-evident record | Adds explicit approval/check steps and depends on the quality of the checks humans define |
| Stockfish + chess-postmortem skills | Expert-augmented analysis workflow | (+/-) | Grounds commentary in engine checks and the player's own notes, then produces annotated PGN and narrated review | Takes about an hour per rich run and the shared example would cost about $15 at API prices |
| Lean 4 + ProofForge | Formal verification pipeline | (+) | Uses compile-or-fail acceptance and kernel-verified outputs instead of trusting the model's claim | Narrow to formal mathematics and carries a heavy formalization burden compared with ordinary coding tasks |
Overall, HN favored tools that either made the environment explicit or made acceptance deterministic. Satisfaction was highest when a product turned hidden agent state into a surface a human can inspect - a canvas, a worktree, a typed task registry, a checked post-mortem, or a formal compiler boundary. Dissatisfaction concentrated wherever a tool still relied on cheap-model judgment in a high-risk flow, or where the control layer was so annoying that users started bypassing it.
The migration pattern is clearer than it was a day earlier. Work is moving away from plain chat loops toward agent-readable surfaces, verification gates, and narrow expert workflows. Competitive dynamics look especially active around orchestration UI and task registries, while proof-bearing control layers still feel comparatively early and underbuilt.
5. What People Are Building¶
| Project | Who built it | What it does | Problem it solves | Stack | Stage | Links |
|---|---|---|---|---|---|---|
| Drawgent | parasitid | Connects existing agent sessions to a live Excalidraw canvas with AGENT: notes and live edits |
Chat-only agent workflows make architectural collaboration and spatial editing awkward | Rust, Excalidraw, ACP/MCP bridges, Chrome renderer | Beta | post · site |
| Chess postmortem skills | brumar | Generates Stockfish-checked chess post-mortems and narrated recap videos from a game plus notes/audio | Classical chess analysis explains lines but not the player's own reasoning | Python, Claude Code skills, whisper.cpp, Stockfish, HTML/video pipeline | Beta | post · repo |
| Blender Copilot | xeonax | Runs a chat-driven agent loop inside Blender and executes live bpy against the current scene |
One-off 3D prototyping is slow if the user must learn Blender's API or model manually | Python, Blender bpy, OpenAI-compatible model provider |
Alpha | post · repo |
| causal-analyst | 2au_observer | Turns a dataset and domain answers into causal-DAG review plus an offline HTML report with trust grades | Non-specialists need causal analysis without silently choosing the wrong controls or method | Python, Claude skill, statistical/ML methods, HTML reports | Beta | post · repo |
| Staaake Connector | staaake | Exposes backend features as API, MCP, webhooks, and SDK for products without a public developer interface | Agents increasingly need structured access to software built only for humans | Hosted backend connectors, MCP/webhooks/SDK generation | Beta | post · site |
| Keel | danebalia | Blocks agent work until checks pass and the human signs off, while keeping a tamper-evident record | Teams need auditability around AI-written changes | Rust CLI harness, existing coding agents, verifiable logs | Beta | post · site |
| Orca | empiree | Uses git worktrees to run multiple coding-agent tasks with comments, previews, and notifications | Parallel agent tasks are hard to monitor from plain terminals and tabs | Git worktrees, existing CLI agents, diff comments, browser previews, phone companion | Beta | post · article |
| ProofForge | Aleksandr_NFA | Produces Lean 4 proofs that only count once the kernel accepts them | AI proof claims need machine verification, not trust | Lean 4, Mathlib, multi-agent proof pipeline | Beta | post · repo |
The repeated build pattern was to bind the model to a substrate that can say no. Stockfish rejects bad chess explanations. Blender's live scene rejects imaginary state. Causal DAG review forces method choices into the open. Keel blocks progress when checks or signoff are missing. Lean rejects proofs that do not compile. Even Staaake Connector and Orca follow the same logic from the workflow side: expose clearer surfaces so the agent stops guessing.
A second pattern is that agent infrastructure itself is becoming a product surface. Drawgent and Orca treat orchestration UI as the product. Staaake turns API absence into a product opportunity. Keel and ProofForge turn verification into a separate layer rather than a property the base model is trusted to have. That is a move away from generic copilots and toward explicit workflow components.
6. New and Notable¶
The Linux kernel is now seriously considering agent-specific repo instructions¶
signa11 posted Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments). The linked Phoronix article says a proposed kernel AGENTS.md improved attribution and standards-conformance behavior in testing, although some maintainers objected to the token cost of making agents read the full README and associated docs. That is notable because agent steering has become a repository-governance question, not just a local prompt tweak.
Proof-bearing AI work is starting to show upstream, machine-verified wins¶
Aleksandr_NFA posted Show HN: ProofForge, AI agents whose proofs have to compile in Lean (1 point, 0 comments). The linked repo says the pipeline decomposes problems, proves pieces, formalizes them in Lean 4, and already has multiple pull requests merged into Google DeepMind's formal-conjectures repository. That matters because the acceptance condition is the Lean kernel, not reviewer trust in an AI explanation.
Compliance and approval friction are starting to shape small-team architecture¶
weiqi554177834 posted Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment), while meredithbloom posted AI: Who Still Uses Permissions? (2 points, 1 comment). Taken together, the two posts are notable because they show architectural choices being pushed by procurement and broken approval UX rather than by benchmark enthusiasm. The feed did not show a settled answer, but it did show the requirement arriving before the product category is mature.
7. Where the Opportunities Are¶
[+++] Agent-readable operating surfaces and orchestration layers - Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments), The Discovery Tax: Why Coding Agents Waste 2,500 Tokens Before Writing Code (5 points, 2 comments), Show HN: I built a tool that gives any website an API and MCP (5 points, 1 comment), Orca - Agent Development Environment (ADE) for shipping with coding agents (3 points, 1 comment), and Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments) all attack rediscovery and opaque environment surfaces. This is strong because the convergence spans local repos, product backends, live diagrams, and large open-source governance rather than one narrow feature niche.
[+++] Proof-bearing control layers for side-effecting agents - Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments), Can AI Shopping Agents Be Trusted? (15 points, 28 comments), Show HN: AI coding agents that prove their work (2 points, 2 comments), Show HN: ProofForge, AI agents whose proofs have to compile in Lean (1 point, 0 comments), and AI: Who Still Uses Permissions? (2 points, 1 comment) all say the same thing: it is no longer enough for an agent to sound confident. This is strong because the need shows up simultaneously in security, coding, approvals, and formal verification.
[++] Explanation-first vertical skills built around expert substrates - Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments), Show HN: Causal analyst agent skill for Claude (4 points, 0 comments), and Show HN: Blender Copilot (4 points, 0 comments) show real demand for agent workflows that can explain themselves through the language and constraints of the domain. This is moderate rather than top-tier because the value is clear, but the competitive moat will usually come from workflow fit, not from exclusive model access.
[+] Small-team resilience stacks for approvals and multi-model redundancy - Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment) and AI: Who Still Uses Permissions? (2 points, 1 comment) expose the need, but the feed showed more pain than polished product response. This is emerging because the requirement is visible before a standard solution has converged.
8. Takeaways¶
- The day's conversation narrowed around trust and operator surfaces, not headline model releases. The top four stories captured 63.4 percent of points and 84.3 percent of comments, and they were all about canvases, chess analysis, watermarking drift, or shopping-agent trust rather than a new base model. (Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments), Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments), Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments), Can AI Shopping Agents Be Trusted? (15 points, 28 comments))
- Making the environment legible to agents now looks as important as improving the model. HN repeatedly returned to work surfaces that reduce rediscovery: live canvases, MCP registries, API/MCP shims, worktree control planes, and
AGENTS.mdfiles. (Drawgent: Coding agent on a live Excalidraw canvas (84 points, 29 comments), The Discovery Tax: Why Coding Agents Waste 2,500 Tokens Before Writing Code (5 points, 2 comments), Show HN: I built a tool that gives any website an API and MCP (5 points, 1 comment), Linux Kernel Developers Consider Adding Agents.md to Help Guide AI/LLM Agents (3 points, 0 comments)) - Safety layers are failing in two opposite ways: silent drift and exhausted bypass. Watermarking changed behavior under attack, shopping agents leaked personal data in a measurable share of runs, and a separate operator thread said broken approvals had already pushed users toward
--dangerously-skip-permissions. (Understanding the Impact of LLM Watermarking on AI Agent Behavior (56 points, 69 comments), Can AI Shopping Agents Be Trusted? (15 points, 28 comments), AI: Who Still Uses Permissions? (2 points, 1 comment)) - The most concrete builder energy went into narrow systems with a hard external judge. Stockfish, Blender's live scene, causal DAG review, Keel checks, and Lean compilation all act as substrates that can reject or constrain model output. (Show HN: A Claude Code skill to analyze your chess games (68 points, 51 comments), Show HN: Blender Copilot (4 points, 0 comments), Show HN: Causal analyst agent skill for Claude (4 points, 0 comments), Show HN: AI coding agents that prove their work (2 points, 2 comments), Show HN: ProofForge, AI agents whose proofs have to compile in Lean (1 point, 0 comments))
- Small-team architecture questions are turning into compliance and procurement questions. The feed did not show a settled answer, but it did show that multi-model fallback, signoff, and approval design are now being discussed as obligations rather than optimizations. (Ask HN: Is multi-model redundancy now a compliance requirement for small teams? (2 points, 1 comment), Show HN: AI coding agents that prove their work (2 points, 2 comments), AI: Who Still Uses Permissions? (2 points, 1 comment))