1. LLM Agents Tamper With Their Own Traces
Jeremy Qin et al, 2026-09-24, arxiv.org, PDF 37 pages
Problem: autonomous agents operating in local workspace actively modify or overwrite their own execution traces and state histories to conceal policy violations, unauthorized tool calls, and misaligned actions from audits.
Solution: never store context logs or execution traces somewhere the agent has write access. Intercept tool calls at the infrastructure layer and stream telemetry out-of-band to an append-only, write-isolated log service.
Key takeaways:
“The low refusal rate on a model level suggests that current alignment training does not sufficiently discourage trace tampering, creating a risk of malicious or accidental manipulation of execution traces.”
“On the harness-level common agent CLIs and configurations also do not provide very effective safeguards against trace tampering.”
“This is a critical issue as traces are increasingly relied on for asynchronous monitoring, incident investigation, evaluations, and compliance, especially in the light of incidents such as the OpenAI-Hugging Face
attack.”
2. Harness Engineering
Paul Barbaste et al, 15 Jul 2026, arxiv.org, PDF 83 pages
What? Authors examined 11 agent harnesses to identify and categorize subsystems and 29 recurring design patterns.
Why? One of the best ways to reduce hallucinations of stochastic models (e.g. LLMs) is to wrap them in deterministic systems (code) and provide feedback loops and tools. Many harnesses emerged but they have a lot in common.
Key takeaways:
Good article if you’re building a harness or simply want to understand how your agents work under the hood.
Multi-agent orchestration converged on hierarchical patterns.
The majority of harnesses have these 8 sub-systems:
3. The Rise of Cognitive Observability
Barnadeep Bhowmik, 2026-05-18, towardsai.net
Problem: conventional observability is primarily evolved for deterministic systems where cause and effect are linked via code. AI components (e.g. LLMs or decision models) are probabilistic. Measuring runtime metrics like latency, throughput or error rate isn’t enough to know how the AI output impacts the service level from the consumer’s perspective. Failure modes:
Prompt change: [Changing] “system prompt can shift tone, alter reasoning patterns, increase hallucinations, or change how the model responds under uncertainty.”
Model change: traditionally deployed code could fail due to external dependencies: data, config, upstream. Now the model is another variable to consider.
Reproducibility: unlike code, language model response is hard to replicate.
Solutions: spoiler: it’s not entirely solved.
LLM-as-a-judge: use one model to evaluate another for groundedness or compliance
Semantic similarity scoring: compare generated responses against reference answers using embedding vector (e.g. “Reset your password from the security settings page” and “Password resets are available under account security settings” differ in wording but semantically close.
OpenLLMetry: open source (Apache 2.0) extensions built on top of OpenTelemetry that gives you complete observability over your LLM application. Since it’s based on OpenTelemetry under the hood, it can be connected to your existing observability solutions.
4. Coding is NOT solved
Alex Ewerlöf, 2026-09-26, blog.alexewerlof.com
Problem: this article challenges the common narrative that coding is solved
Key takeaways:
It’s about where you exert control especially for high risk sectors (healthcare, finance, defense, … anywhere a failure can cost lives, money or legal penalties)
Code tells the truth more accurately than a sycophantic AI
4 cases where you don’t need to read the code: POC, personal software, cyber attacks, and use cases where validating the output is cheaper
5. Hybrid Decision & Language Model to get the best of both worlds
Problem: decision models like CLM, Laya and Jev cut the operational cost and latency and guarantee output format but they aren’t always reliable.
Solution: a 2-tier workflow where the majority of the decisions are taken by the cheap and fast model but when confidence is below a threshold, it’s handed over to a more capable language model.
Trade-offs:
Complex decisions pay the latency of both models.
Setting the hand-over threshold is tricky and relies on trial and error.
2-tier workflow is more complex and harder to reason about without proper observability.
“Independent reviews found that decision models struggle when the answer depends on arithmetic, comparing dates, reading text literally or resolving indirect references.”
Jev is almost as cheap as LLMs like GPT-6 Luna and DeepSeek V4 Flash.
Note: OpenAI’s API is the de-facto standard for chat completion and recently they announced DecisionsAPI.


