Field notes / agentic AI

What is actually changing, and what it changes about the way you build.

The field moves faster than any curriculum can be rewritten. So this is where the reading gets published, rather than just claimed. Every finding below is dated, carries its original source, and is read by a person before it appears here. A preprint is labelled a preprint, and a vendor's research is labelled a vendor's.

7 current findings · last updated 15 August 2026 · nothing older than 90 days

Sourced from
01 / 2 findings

Reliability & evaluation

Where agentic systems fail in ways a token count or an eyeballed summary will not show you.

Compressing system-side control context has a sharp, non-linear reliability cliff

"Control Under Compression" (arXiv 2608.01056, 2 Aug 2026) looks at agent control contexts rather than conversation history. A control context is the persistent system-side instruction set. It specifies tools, arguments, policies, execution protocols and recovery. The benchmark ran 15,525 environment-verified runs. Retaining 75% of the context held success near the full-context baseline, at 92.7% and 92.4% against 93.8%. Between 50% and 35% the methods diverged sharply. At 35% they scored 47.0%, 39.0% and 19.9%, depending on the compression method. Failures appeared mainly as tool-execution and action-parsing errors. Reliability varied enough between control contexts that the authors argue no compressor ranking holds universally, and that each context needs its own qualification. So trimming tool and policy prompts is a runtime-reliability decision, and it needs executable testing per context. The safe-looking zone ends abruptly.

arxiv.org agentic-aitool-designcontext-managementreliability

Context compaction shows up as a reliability failure, not just a cost lever

A preliminary empirical study (Min et al., arXiv 2608.06503, 6 Aug 2026) looks at agents that repeatedly compress their own context. It finds that this compression weakens the influence of recent interactions. The agent then blocks more actions, explores the same ground again, and behaves differently from run to run. Token-count metrics never show any of this. The authors propose TRACE. It evaluates each individual compaction event by running paired closed-loop continuations from the same environment state, then optimises the compression prompt while leaving all models frozen. If your agent compacts, treat each compaction boundary as a discrete event and test it by what the agent then does, rather than by reading the summary and judging it by eye.

arxiv.org agentic-aicontext-managementreliabilityevaluation
02 / 3 findings

Security & containment

The attack surface a system gains the moment it acts on content it did not author.

Early academic work argues the agent threat model is broader than instruction injection

A preprint led by Seoul National University argues that agent security research has been too narrow. Prior work has focused mainly on indirect prompt injection. The most-studied category within that is instruction injection, where attacker-controlled untrusted data is read as an instruction. The paper argues that data injection is a realistic threat class in its own right. Data injection means poisoning what the agent believes to be true, rather than what it is told to do. If your agent threat model is built only around "untrusted text might contain commands", this is a reason to widen it.

arxiv.org securitythreat-modellingagentic-airesearch

Indirect prompt injection against web-browsing agents documented in live campaigns

Zscaler ThreatLabz documented real-world campaigns that embed instructions in web content aimed at AI agents. They describe indirect prompt injection as an attack that hides malicious instructions inside content an AI agent retrieves, such as websites, documents or email, in order to influence the agent's reasoning while it works. Their test setup matters. They built an autonomous agent with web browsing and payment-execution tools. They ran it fully sandboxed, with no real funds. They deliberately configured it with no spending limits, in order to measure the largest possible exploitation surface. The engineering takeaway is that spend caps, tool-scope limits, and human confirmation on actions you cannot undo are the controls doing the actual work. Model-level instruction hardening is not.

zscaler.com securityprompt-injectionagentic-aicontainment

Statelessness moves MCP authorization to the application layer, and most MCP risk was never in the transport

Removing the session also removes the per-session authorization context. As Google's write-up puts it, "As the responsibility of managing state shifts from the transport layer to the application layer, security becomes paramount." An independent security review of the release is blunter. Statelessness changes where authorization happens and how you deploy the server. But the vulnerabilities found in MCP servers sit in the functionality those servers expose, and the specification does not touch that. The spec upgrade buys scale and cheaper operations, not safety. Every request must now carry and re-verify its own authorization, and tool-level authorization design remains entirely your problem.

equixly.com mcpsecurityauthorizationarchitecture
03 / 2 findings

Protocol, tooling & architecture

Releases that change how agentic systems get built — and what they make legacy.

MCP migration has sharp edges: a silent version downgrade and three deprecated capabilities

Beta releases of the Python, TypeScript, Go and C# SDKs are available with support for the 2026-07-28 spec. Here is the trap. The streamable HTTP transport accepts 2026-07-28 only when you set StreamableHTTPOptions.Stateless = true. Leave it unset and clients negotiate down to 2025-11-25. That is a silent change in behaviour rather than an error. So add an explicit assertion on the negotiated protocol version to your integration tests. The other breaking changes are confined to the capabilities the specification deprecates: roots, sampling and logging. If your agent design depends on server-initiated sampling, that is now on a deprecation path. Note also that server/discover is optional for clients to call, but mandatory for every 2026-07-28 server to implement.

blog.modelcontextprotocol.io mcptooling-releasemigration

MCP went stateless: the 2026-07-28 spec is a breaking change for remote MCP server deployment

The 2026-07-28 Model Context Protocol specification shipped with a stateless protocol core. It also brought Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework, and updated Tier 1 SDKs. This is the largest revision of the protocol since launch. The initialize/initialized handshake (SEP-2575) and the Mcp-Session-Id header (SEP-2567) are removed entirely. Every request is now self-describing. The protocol version, client info and capabilities travel in a _meta field inline on every request. A remote MCP server used to need sticky sessions, a shared session store, and deep packet inspection at the gateway. It can now run behind a plain round-robin load balancer, route traffic on an Mcp-Method header, and let clients cache tools/list responses for as long as the server's ttlMs permits. If you have an MCP server in production behind a Redis session store, that architecture is now legacy. Plan the migration rather than inheriting it.

blog.modelcontextprotocol.io mcptooling-releasearchitecture
Method / how this page is made

An agent gathers it. A human decides what you see.

This page is itself a small piece of the craft it teaches. It is a retrieval agent with a bounded budget and a human review gate. The agent is deliberately not allowed to say anything it cannot source.

  • 01

    One research pass per topic, on a schedule

    The retriever runs weekly across a fixed set of topics. Those are architecture patterns, evaluation and reliability, agent security, and releases that change how systems get built. Each topic gets its own budget. A shared budget lets the first topic use it all up, and the agent then reports thin findings instead of admitting the gap.

  • 02

    Nothing publishes itself

    The scheduled run opens a pull request rather than committing straight away. Sunil reads what the agent gathered before a prospect does. Every item carries a private note on source quality and on what could not be confirmed. That is why the wording here is careful about the difference between an unreplicated preprint and a settled result.

  • 03

    The agent cannot write our own numbers

    Prices, dates, and seat counts are deliberately out of scope for the retriever, and they never appear on this page. They live in one place and are published from there. So a figure can never be current on an offer page and out of date in a finding.

  • 04

    Findings expire

    Anything older than 90 days drops off automatically. A "latest" page quietly serving year-old news is worse than no page at all.

This is the reading. The programme is the judgement.

Findings like these are raw material. Knowing that context compaction is a reliability problem is not the same as knowing what to do about it in a system you are accountable for. That is what the programme is for.

Or ask the site's agent about any finding above. It answers from the same notes.