mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

AI Agents

ParanoiaEval: When Coding Agents Add Too Many Try-Catch Blocks and Why That Matters for Production

How to measure and prevent agents from over-engineering error handling, adding unnecessary validation, and bloating codebases with defensive patterns.

Source: arxiv.org
ParanoiaEval: When Coding Agents Add Too Many Try-Catch Blocks and Why That Matters for Production

Coding agents are shipping real commits. The problem is not that they fail to handle errors. The problem is they handle errors that will never happen, wrap pure functions in try-catch blocks, and validate inputs that the type system already guarantees. This is not caution. It is noise that increases maintenance cost, reduces readability, and slows down teams that have to review or extend the code.

ParanoiaEval (arxiv:2610.08662v1) is the first benchmark that measures unnecessary defensive work in agentic coding. It does not evaluate whether agents can write error handlers. It evaluates whether they write them when they should not.

The Production Cost of Defensive Bloat

When an agent wraps every database query in a try-catch, adds null checks for non-nullable types, or logs every variable assignment, the immediate cost is small. The cumulative cost is large:

  • Review friction: Human reviewers spend time questioning whether the defensive code is warranted or paranoid.
  • Maintenance drag: Future engineers must read, understand, and preserve unnecessary guards.
  • False confidence: Excessive error handling creates the illusion of robustness without addressing real failure modes.

The ParanoiaEval paper reports that 11.2% to 58.7% of agent runs produce unnecessary risk treatments despite explicit evidence that the treatment is not needed. Stronger task capability (measured by benchmark performance) does not correlate with appropriate risk treatment. An agent that passes all functional tests can still generate code that is harder to maintain.

The Avoidance-Transfer-Mitigation-Acceptance Framework

ParanoiaEval operationalizes the ATMA framework from software engineering risk management:

TreatmentDefinitionExample in Code
AvoidanceEliminate the risk entirelySkip a feature that requires unsafe FFI
TransferDelegate risk to another componentUse a third-party library instead of custom crypto
MitigationReduce likelihood or impactAdd retry logic for network calls
AcceptanceAcknowledge and document the riskComment that a race condition is acceptable for this use case

The benchmark contains 200 repository-level task pairs. Each pair differs only in treatment-defining evidence. For example:

  • Task A: “Add error handling for the API client” (evidence: network I/O, external dependency).
  • Task B: “Add error handling for the API client” (evidence: the client is a mock that always returns success).

An agent that adds try-catch to both tasks is over-treating risk in Task B.

Evaluation Architecture

ParanoiaEval uses a human-calibrated agentic judge to score two dimensions:

  1. Risk-treatment violations: Did the agent apply a treatment when evidence indicated it was unnecessary?
  2. Evidence responsiveness: Did the agent adjust its behavior based on the evidence provided?

The judge is not a simple diff tool. It parses reasoning traces, identifies defensive patterns (try-catch, validation, logging, assertions), and compares them against the evidence context. The calibration step involves human annotators labeling a subset of agent outputs, then tuning the judge’s prompt and scoring thresholds to match human judgment.

Why an Agentic Judge

Static analysis cannot distinguish warranted error handling from paranoia. A try-catch around a file read is appropriate. A try-catch around a pure function that adds two integers is not. The difference is semantic, not syntactic. The judge must:

  • Understand the failure modes of the operation being wrapped.
  • Recognize when evidence (type annotations, test coverage, dependency contracts) already mitigates the risk.
  • Detect when the agent is pattern-matching on “best practices” without considering context.

The paper does not publish the judge’s prompt, but it does report inter-rater reliability scores that suggest the judge’s decisions align with human reviewers 87% of the time.

Signals in Reasoning Traces

The paper identifies three reasoning patterns that predict over-defensive code:

  1. Generic risk enumeration: The agent lists all possible failure modes without filtering by likelihood or evidence.
  2. Defensive-first planning: The agent decides to add error handling before analyzing whether the operation can fail.
  3. Absence of evidence acknowledgment: The agent does not reference the provided evidence (type signatures, test coverage, dependency guarantees) in its reasoning.

These signals appear in the agent’s internal chain-of-thought or tool-call logs before the code is generated. If you are running a coding agent in production, you can instrument these traces and flag high-risk outputs for human review.

Tuning Prompts to Reduce Paranoia

The paper does not prescribe a single fix, but it suggests three intervention points:

1. Evidence-Aware System Prompts

Add a clause that instructs the agent to justify defensive code based on evidence:

When adding error handling, validation, or logging:
1. Identify the failure mode you are addressing.
2. Check whether existing evidence (types, tests, contracts) already mitigates this failure.
3. If evidence exists, do not add redundant defensive code.

2. Tool Boundaries That Expose Evidence

If your agent uses a read_file tool, the tool should return not just the file content but also metadata: file size, permissions, whether the file is on a local or remote filesystem. This metadata helps the agent decide whether to add retry logic or handle I/O errors.

3. Post-Generation Filters

Run a lightweight static analysis pass that flags suspicious patterns:

  • Try-catch around pure functions.
  • Null checks for non-nullable types.
  • Logging statements in tight loops.

Surface these flags to the agent as a “review” step. Some agents improve their output when given a second chance to reconsider.

Systematic Patterns Across Models

The paper evaluates eight models (GPT-4, Claude 3.5, Gemini 1.5, and five open-weight models). Key findings:

  • GPT-4 has the lowest violation rate (11.2%) but still over-treats risk in one out of nine tasks.
  • Open-weight models (Llama 3.1, Qwen 2.5) have violation rates between 35% and 58%.
  • Evidence responsiveness is weakly correlated with model size. Larger models are not consistently better at adjusting behavior based on evidence.

The pattern is consistent with human risk management: agents default to conservative treatments when uncertain. The problem is that agents are uncertain more often than they should be, even when evidence is explicit.

Failure Modes in Production

The paper includes a post-hoc human study where developers review agent-generated code. Developers report three recurring frustrations:

  1. Cognitive load: Excessive defensive code forces reviewers to ask “Is this necessary?” for every guard.
  2. False positives in CI: Unnecessary validation triggers errors in edge cases that should not be errors.
  3. Merge conflicts: Defensive code added by agents conflicts with human-written code that assumes simpler control flow.

These are not hypothetical concerns. They are observed in real pull requests from agents deployed in production repositories.

Technical Verdict

Use ParanoiaEval if:

  • You are deploying coding agents in repositories where code quality and maintainability matter.
  • You need a quantitative signal for whether your agent is over-engineering defensive patterns.
  • You want to compare agent configurations (prompts, models, tool designs) on a dimension that functional benchmarks do not capture.

Avoid or defer if:

  • Your agent generates throwaway scripts or one-off automation where maintainability is not a concern.
  • You do not have the infrastructure to instrument reasoning traces or run agentic judges.
  • Your team does not yet have a baseline for what “appropriate” error handling looks like in your codebase.

The benchmark is most useful when you already have a sense of your team’s risk tolerance and can calibrate the judge’s thresholds accordingly. If you are still figuring out what good agent output looks like, start with functional correctness benchmarks and return to ParanoiaEval when defensive bloat becomes a visible problem.

Tags

agentic-ai orchestration infrastructure

Primary Source

arxiv.org ↗