Skip to content
Victor Del Puerto
Go back

Functional Scars: Turning Repeated AI Agent Mistakes Into Runtime Controls

My coding agents kept repeating mistakes I had already corrected, days after I explained them, with the explanation sitting in their memory files.

On 17 September an agent added a line to a notes index in my workspace. The line was supposed to say transcribe with base on CPU. It said transcribe with on CPU. The same thing in a terminal:

$ echo "transcribe with `base` on CPU"
/usr/bin/bash: line 1: base: command not found
transcribe with  on CPU

Inside double quotes, bash runs whatever sits between backticks and pastes the output in its place, as the POSIX shell specification says it should. Agents write Markdown, Markdown puts every file name in backticks, and the exit code stays 0. That evening it happened twice more within two hours. The third time, a backticked name matched a project file, and bash ran the file as a shell script. The trap had been in the agent’s memory notes since 2 September.

I run DG Ingeniería, a small engineering firm in Paraguay. Claude Code and Codex write and patch files in my workspace every day, and I keep a record of every correction I give them. Out of that record I built a tool I call functional scars. This article covers why it exists, how it works, three cases measured in my own workspace, and what it does not solve.

Why it exists

A correction can fail at several points before it changes anything. It has to be recorded, found again when a similar task comes up, read by the agent, and applied to the action. Memory systems and instruction files cover the first steps. The last ones belong to the model, and a note competes with a habit that usually works: of the 3,885 heredoc writes to Python files in my transcripts before one of the scars below, 3,590 passed through the shell intact.

I think of it in four layers:

Layers 1 to 3, the instruction file, the knowledge base and the memory store, all feed the model, which reads them and decides. Layer 4 is a check outside the model that sits between the action the model proposes and its execution, and runs every time.

Figure 1. Three layers hand the model information. The fourth sits on the path of the action.

The first three hand the agent information. The fourth acts on the action itself. Engineers learn the same way: a junior reads the textbook, and a senior carries the weight of the mistakes that went wrong in production, which changes the next decision before any deliberation starts. A functional scar is that weight, written as code.

Others have reached a similar place from other directions. Dmitriy Fedoryshchev describes a demotion ladder, where a rule that keeps failing moves from prose to something harder to skip. Bold text in an instruction file does not buy attention, so making the prose louder is no substitute. In research, AgentSpec specifies runtime enforcement as triggers, predicates and actions, and TRACE compiles a user’s chat corrections into checks that run before a coding agent finishes.

What a functional scar is

A functional scar is a versioned check, derived from a documented failure, that runs outside the model at a named event of the agent runtime. I write each one as a contract with six fields, so it can be tested, measured and retired later. For the backtick case:

The six-field contract of the backtick scar. Origin: traces since 31 July; 21 bash errors in the 17 days the note sat in memory. Applies when: a Bash command puts an unescaped backtick inside double quotes. Obligation: the text that reaches the file is the text the agent wrote. Intervention: before the Bash call runs, deny it, say why, name the alternatives. Evidence: one log line per firing with time, session, scar and version, action, payload hash. Review: a test suite of real command shapes; every miss becomes a case.

Figure 2. The contract of the backtick scar.

The intervention can take three forms, and they promise different things. Context added before the model answers can still be ignored. A denial before a tool call runs means the call does not happen. A check after the call reports what reached the disk.

How it works

Claude Code and Codex both run hooks: commands the host executes at named events of the agent’s loop, with the event passed as JSON on standard input (Claude Code, Codex). A scar lives in one of those hooks.

One turn of the agent: the prompt is submitted, the agent proposes a tool call, the tool runs, and the result comes back. Scars can act at three events. When the prompt is submitted they add context, which the model can still ignore. Before the tool call they can deny it. After it they check what reached the disk. Before a tool call, the host sends the event as JSON to fscars, which loads the project's scars and asks each one whether it matches. If none does, the call runs. If one does, fscars answers in the host's format with a denial and the reason, and writes one line to the log.

Figure 3. Where scars act in one turn of the agent, and what happens when one fires before a tool call.

fscars is the open-source Python package I maintain for this. It registers one entry point, python -m fscars.run_hook, for every event, loads the project’s scars from .fscars/scars/, asks each one whether it applies, answers in the format the host expects, and writes one line per firing to .fscars/logs/fires.jsonl. A scar is a small Python class that says when it applies, what it checks and what it answers:

class BacktickInDoubleQuotes(FunctionalScar):
    scar_id = "backtick-in-double-quotes"
    name = "Backticks inside double quotes run as commands"
    rule = ("Text with backticks never travels inside a double-quoted shell argument: "
            "write it to a file, use single quotes, or use $(...) when substitution is intended.")
    severity = Severity.BLOCK
    event_type = HookEventType.PRE_TOOL_USE
    tool_matchers = ("Bash",)

    def matches(self, payload: HookPayload) -> bool:
        return bool(backticks_in_double_quotes(payload.tool_input.get("command") or ""))

    def build_output(self, payload: HookPayload) -> ScarOutput:
        hits = backticks_in_double_quotes(payload.tool_input.get("command") or "")
        return ScarOutput(
            additional_context=f"[{self.scar_id}] bash would run {', '.join(hits[:3])} as a command. {self.rule}",
            system_message=f"{self.scar_id}: blocked",
            block=True,
        )


scar = BacktickInDoubleQuotes()

The hard part is backticks_in_double_quotes, a short parser that follows bash’s quoting. The full file, and how I tested it against real bash and both hosts, are in the appendix at the end.

Three cases, measured

Each case below comes from my own workspace, measured on Claude Code session transcripts, which record every command exactly as the agent wrote it, together with its result.

Case 1: Python files written through a heredoc

On Windows, Claude Code’s Bash tool halves runs of backslashes before bash reads the command, except runs right before a double quote (open issue). Agents patching .py files through shell heredocs kept breaking string literals. With a note in memory, the breakage continued. After a scar that denies heredocs writing .py files, the breakage traced to them went to zero.

Across 102,298 Bash calls between 17 July and 4 October, heredoc writes to .py files fell from 6.7% of Bash calls before the scar to 0.8% after. The share of sessions that tried at least once did not fall; the repetition inside a session did, from 31.6 attempts per session to 3.3. Before the scar, 47 unterminated string literal errors traced back to those heredocs. After it, the count was 0, although only four of those heredocs ran at all.

Before and after the heredoc scar. Heredoc writes to .py files fell from 6.7% to 0.8% of Bash calls. Attempts per session, among sessions that tried, fell from 31.6 to 3.3. Unterminated string literal errors traced to those heredocs fell from 47 to 0.

Figure 4. Case 1 before and after the scar, over 102,298 Bash calls from 17 July to 4 October.

The cost is part of the result. 92% of the blocked commands carried text the transport would have left intact. By the failure rates of the first window, the blocks themselves prevented an estimated four to six errors over 24 days, part of which a sibling scar would also have caught. The broad rule targets a habit, pushing code and text through the shell, and the same habit produced two more scars in the following eight days, the backtick case among them; the price of that choice is measured. The measurement script for this case went through two rounds of review by a second agent, and the first round corrected an error that had made the effect look larger than it was.

Case 2: backticks inside double quotes

This is the case from the opening. Across the same transcripts I counted the Bash commands that put an unescaped backtick inside double quotes, using the same parser as the scar above, and checked the ones that ran for an error from bash in their output. An error is only the visible part: a backticked word that happens to be a real command, or an empty pair, substitutes without a sound, so the bash-error bars are a lower bound.

Case 2 by window, per thousand Bash calls. Before the note, to 2 September: 49,315 calls, 13 exposures, 13 ran, 9 bash errors. Note in memory, 2 to 19 September: 19,754 calls, 25 exposures, 23 ran, 21 errors. Rule written, no hook, 19 to 21 September: 3,272 calls, 4 exposures, 4 ran, 4 errors. Hook v1, python -c only, 21 to 22 September: 3,428 calls, 3 exposures, 1 ran, 1 error. Hook v2, adding node -e and git commit -m: 26,529 calls, 6 exposures, 1 ran, 1 error.

Figure 5. Case 2 per thousand Bash calls in each window, with the count on each bar. The rule and hook v1 windows last two days and one day, so their rates rest on a handful of commands.

The first trace is from 31 July. The trap went into the agent’s memory notes on 2 September, and the note did not lower the rate: 0.26 exposures per thousand Bash calls before it, 1.27 in the 17 days after, with 21 errors. The comparison is observational, and the workload changed too. The written rule of 19 September changed nothing measurable either: four exposures in two days, four errors.

The first hook watched only python -c. It blocked two commands and missed a git commit -m whose message lost a word. The second version added node -e and git commit -m. After it, the five commands in watched forms were all blocked, and one in another form ran: a grep for Markdown code fences whose backticks swallowed half the command, so the first search never ran and the second printed every line of its file. That miss is why the scar in this article checks the mechanism in any command, without a list of commands.

The deployed hook also blocks a dollar sign inside double quotes, which bash expands as well. Of the 102 denials in the transcripts, 95 were for a dollar sign and 7 for backticks. I have not labeled how many of those dollar signs were meant to expand.

Case 3: a memory hook that can only remind

Some obligations have no tool call to hold. One scar in my workspace matches each prompt against an index of memory notes and injects the names of the closest ones, so the agent opens them before answering. In a blind census of its firings, 699 hash-verified firings over two months, a single model judge found a relevant note in 471, and the agent opened that note later in the session after 111 of them, 23.6%, an upper bound.

Case 3 as a funnel. 699 hash-verified firings of the memory hook. In 471 of them a model judge found a relevant note. After 111 of those, 23.6%, the agent opened the note later in the session.

Figure 6. Case 3: the reminder arrived every time, and the relevant note was opened after fewer than one in four.

The hook delivered its reminder every time. Three times out of four, the relevant note stayed closed. Where the obligation shows up in a command or a file, a scar can act on it. Where it lives in what the model reads, a scar can only remind.

What it solves, and what it does not

A functional scar solves a specific problem: a mistake that already happened, that has an observable signature, and that shows up at a moment the host exposes, a prompt, a tool call or a file on disk. For that kind of mistake it turns “please remember” into a check that runs every time, leaves a record, and can be measured.

It does not catch a new kind of mistake, and it cannot enforce a judgment call that has no observable signature. It covers only the events the host hands to it. Its trigger can miss forms of the same mistake, as Case 2 shows, and every firing on a harmless command costs a retry, as Case 1 shows. It changes no model weights. And a scar is code that runs on every tool call, so it fails like code: while measuring for this article I found that the pattern the deployed backtick scar uses for git commit -m backtracked exponentially on commands with many flags, 15 seconds at 24 flags, until I removed a redundant alternative. A scar also needs an end: the backslash scar should retire the day the upstream bug is fixed.

What none of this proves

All of it comes from one workspace and one operator. The before-and-after comparisons are observational: models, workloads and the rules’ own text changed at the same time, and nothing was randomized. The tests show that the code does what the tests specify, and nothing more. What would settle the question is a controlled comparison of the same correction stored as a note, delivered as context, and enforced as a scar, on matched tasks with an outside check of the result. The preprint describes that design and the evidence so far.

If you want to try it

pip install fscars
cd your-project
fscar init

fscar init creates .fscars/, scaffolds six starter scars and registers the hook for Claude Code; fscar init --adapter codex registers it for Codex. Put the file from the appendix below in .fscars/scars/, and fscar log will show every time it fires.

Appendix: the full scar

"""Scar: a Bash command puts an unescaped backtick inside double quotes, so bash runs it before the program sees the text."""
import re

from fscars.core.fire import Severity
from fscars.core.payload import HookEventType, HookPayload
from fscars.core.scar import FunctionalScar, ScarOutput

HEREDOC = re.compile(r"<<-?\s*(['\"]?)(\w+)\1")


def backticks_in_double_quotes(cmd):
    """Contents of unescaped backticks that bash would run from inside double quotes."""
    lines, kept, i = cmd.split("\n"), [], 0
    while i < len(lines):  # a heredoc body with a quoted delimiter is literal: drop it
        kept.append(lines[i])
        opener = HEREDOC.search(lines[i])
        i += 1
        if opener:
            body = []
            while i < len(lines) and lines[i].strip() != opener.group(2):
                body.append(lines[i])
                i += 1
            i += 1
            if not opener.group(1):  # unquoted delimiter: the body expands like double-quoted text
                kept.append('"' + "\n".join(body).replace('"', "") + '"')
    text, found, stack, i = "\n".join(kept), [], ["plain"], 0
    while i < len(text):
        top, c = stack[-1], text[i]
        if top == "single":
            if c == "'":
                stack.pop()
            i += 1
        elif c == "\\":
            i += 2  # an escaped character never opens or closes anything
        elif top == "double" and c == "`":
            end = text.find("`", i + 1)
            end = len(text) if end == -1 else end
            found.append(text[i + 1:end])
            i = end + 1
        elif text.startswith("$(", i):
            stack.append("subshell")  # quoting starts over inside $( )
            i += 2
        elif top == "double":
            if c == '"':
                stack.pop()
            i += 1
        else:  # plain text, or inside $( )
            if c == "#" and (i == 0 or text[i - 1] in " \t\n;&|("):
                end = text.find("\n", i)
                i = len(text) if end == -1 else end  # a comment: an apostrophe here quotes nothing
                continue
            if c == "'":
                stack.append("single")
            elif c == '"':
                stack.append("double")
            elif c == ")" and top == "subshell":
                stack.pop()
            i += 1
    return found


class BacktickInDoubleQuotes(FunctionalScar):
    scar_id = "backtick-in-double-quotes"
    name = "Backticks inside double quotes run as commands"
    rule = ("Text with backticks never travels inside a double-quoted shell argument: "
            "write it to a file, use single quotes, or use $(...) when substitution is intended.")
    severity = Severity.BLOCK
    event_type = HookEventType.PRE_TOOL_USE
    tool_matchers = ("Bash",)

    def matches(self, payload: HookPayload) -> bool:
        return bool(backticks_in_double_quotes(payload.tool_input.get("command") or ""))

    def build_output(self, payload: HookPayload) -> ScarOutput:
        hits = backticks_in_double_quotes(payload.tool_input.get("command") or "")
        return ScarOutput(
            additional_context=f"[{self.scar_id}] bash would run {', '.join(hits[:3])} as a command. {self.rule}",
            system_message=f"{self.scar_id}: blocked",
            block=True,
        )


scar = BacktickInDoubleQuotes()

The parser follows bash’s quoting: single quotes are literal, a backslash escapes the next character, quoting starts over inside $(...), a comment quotes nothing, and a heredoc with a quoted delimiter is literal. I checked each test case against real bash first, with a harmless echo INJ between the backticks, and then ran the file through the real entry point with the Claude Code adapter and the Codex adapter.

Both adapters blocked, with exit code 2, a git commit -m and a python -c whose double-quoted text carried backticks, and two commands where bash does run the backticks even though an apostrophe appears earlier, once inside a heredoc body and once inside a comment. Both let through the same commit message in single quotes or with escaped backticks, echo "today is $(date +%F)", a backtick in single quotes inside $(...), and the commit form Claude Code itself proposes, git commit -m "$(cat <<'EOF' ... EOF)", with backticks in the message. Every block left a line in the log. Three mutants, one ignoring escapes, one without the $(...) context and one that does not skip comments, each fail the test, which is how I know the test can fail.

AI disclosure: I used AI assistance to put this text into English and for editorial and drafting support.


Share this post on:

Next Post
How we use engines with skills and knowledge graphs