Coding agents are shipping real commits. The problem is not that they fail to handle errors. The problem is they handle errors that will never happen, wrap pure functions in try-catch blocks, and validate inputs that the type system already guarantees. This is not caution. It is noise that increases maintenance cost, reduces readability, and slows down teams that have to review or extend the code.
ParanoiaEval (arxiv:2610.08662v1) is the first benchmark that measures unnecessary defensive work in agentic coding. It does not evaluate whether agents can write error handlers. It evaluates whether they write them when they should not.
The Production Cost of Defensive Bloat
When an agent wraps every database query in a try-catch, adds null checks for non-nullable types, or logs every variable assignment, the immediate cost is small. The cumulative cost is large:
- Review friction: Human reviewers spend time questioning whether the defensive code is warranted or paranoid.
- Maintenance drag: Future engineers must read, understand, and preserve unnecessary guards.
- False confidence: Excessive error handling creates the illusion of robustness without addressing real failure modes.
The ParanoiaEval paper reports that 11.2% to 58.7% of agent runs produce unnecessary risk treatments despite explicit evidence that the treatment is not needed. Stronger task capability (measured by benchmark performance) does not correlate with appropriate risk treatment. An agent that passes all functional tests can still generate code that is harder to maintain.
The Avoidance-Transfer-Mitigation-Acceptance Framework
ParanoiaEval operationalizes the ATMA framework from software engineering risk management:
| Treatment | Definition | Example in Code |
|---|---|---|
| Avoidance | Eliminate the risk entirely | Skip a feature that requires unsafe FFI |
| Transfer | Delegate risk to another component | Use a third-party library instead of custom crypto |
| Mitigation | Reduce likelihood or impact | Add retry logic for network calls |
| Acceptance | Acknowledge and document the risk | Comment that a race condition is acceptable for this use case |
The benchmark contains 200 repository-level task pairs. Each pair differs only in treatment-defining evidence. For example:
- Task A: “Add error handling for the API client” (evidence: network I/O, external dependency).
- Task B: “Add error handling for the API client” (evidence: the client is a mock that always returns success).
An agent that adds try-catch to both tasks is over-treating risk in Task B.
Evaluation Architecture
ParanoiaEval uses a human-calibrated agentic judge to score two dimensions:
- Risk-treatment violations: Did the agent apply a treatment when evidence indicated it was unnecessary?
- Evidence responsiveness: Did the agent adjust its behavior based on the evidence provided?
The judge is not a simple diff tool. It parses reasoning traces, identifies defensive patterns (try-catch, validation, logging, assertions), and compares them against the evidence context. The calibration step involves human annotators labeling a subset of agent outputs, then tuning the judge’s prompt and scoring thresholds to match human judgment.
Why an Agentic Judge
Static analysis cannot distinguish warranted error handling from paranoia. A try-catch around a file read is appropriate. A try-catch around a pure function that adds two integers is not. The difference is semantic, not syntactic. The judge must:
- Understand the failure modes of the operation being wrapped.
- Recognize when evidence (type annotations, test coverage, dependency contracts) already mitigates the risk.
- Detect when the agent is pattern-matching on “best practices” without considering context.
The paper does not publish the judge’s prompt, but it does report inter-rater reliability scores that suggest the judge’s decisions align with human reviewers 87% of the time.
Signals in Reasoning Traces
The paper identifies three reasoning patterns that predict over-defensive code:
- Generic risk enumeration: The agent lists all possible failure modes without filtering by likelihood or evidence.
- Defensive-first planning: The agent decides to add error handling before analyzing whether the operation can fail.
- Absence of evidence acknowledgment: The agent does not reference the provided evidence (type signatures, test coverage, dependency guarantees) in its reasoning.
These signals appear in the agent’s internal chain-of-thought or tool-call logs before the code is generated. If you are running a coding agent in production, you can instrument these traces and flag high-risk outputs for human review.
Tuning Prompts to Reduce Paranoia
The paper does not prescribe a single fix, but it suggests three intervention points:
1. Evidence-Aware System Prompts
Add a clause that instructs the agent to justify defensive code based on evidence:
When adding error handling, validation, or logging:
1. Identify the failure mode you are addressing.
2. Check whether existing evidence (types, tests, contracts) already mitigates this failure.
3. If evidence exists, do not add redundant defensive code.
2. Tool Boundaries That Expose Evidence
If your agent uses a read_file tool, the tool should return not just the file content but also metadata: file size, permissions, whether the file is on a local or remote filesystem. This metadata helps the agent decide whether to add retry logic or handle I/O errors.
3. Post-Generation Filters
Run a lightweight static analysis pass that flags suspicious patterns:
- Try-catch around pure functions.
- Null checks for non-nullable types.
- Logging statements in tight loops.
Surface these flags to the agent as a “review” step. Some agents improve their output when given a second chance to reconsider.
Systematic Patterns Across Models
The paper evaluates eight models (GPT-4, Claude 3.5, Gemini 1.5, and five open-weight models). Key findings:
- GPT-4 has the lowest violation rate (11.2%) but still over-treats risk in one out of nine tasks.
- Open-weight models (Llama 3.1, Qwen 2.5) have violation rates between 35% and 58%.
- Evidence responsiveness is weakly correlated with model size. Larger models are not consistently better at adjusting behavior based on evidence.
The pattern is consistent with human risk management: agents default to conservative treatments when uncertain. The problem is that agents are uncertain more often than they should be, even when evidence is explicit.
Failure Modes in Production
The paper includes a post-hoc human study where developers review agent-generated code. Developers report three recurring frustrations:
- Cognitive load: Excessive defensive code forces reviewers to ask “Is this necessary?” for every guard.
- False positives in CI: Unnecessary validation triggers errors in edge cases that should not be errors.
- Merge conflicts: Defensive code added by agents conflicts with human-written code that assumes simpler control flow.
These are not hypothetical concerns. They are observed in real pull requests from agents deployed in production repositories.
Technical Verdict
Use ParanoiaEval if:
- You are deploying coding agents in repositories where code quality and maintainability matter.
- You need a quantitative signal for whether your agent is over-engineering defensive patterns.
- You want to compare agent configurations (prompts, models, tool designs) on a dimension that functional benchmarks do not capture.
Avoid or defer if:
- Your agent generates throwaway scripts or one-off automation where maintainability is not a concern.
- You do not have the infrastructure to instrument reasoning traces or run agentic judges.
- Your team does not yet have a baseline for what “appropriate” error handling looks like in your codebase.
The benchmark is most useful when you already have a sense of your team’s risk tolerance and can calibrate the judge’s thresholds accordingly. If you are still figuring out what good agent output looks like, start with functional correctness benchmarks and return to ParanoiaEval when defensive bloat becomes a visible problem.