Abstract
Background. AI coding agents can build working software from a few sentences, but they start every session without the project's facts, the preferences of the person they work for, or the lessons of earlier sessions. Giving that work a lasting structure is left to the developer, and the frameworks that try usually ship large rule sets that nobody tests.
Case. City App Framework set out to give AI-assisted software and web development that structure: a project scaffolder (v1), a "city" of AI roles (v2), and a rulebook of about 40 convention documents on how to build (v3).
Method. We tested which parts of that structure change what an agent does: 285 fresh, headless Claude Code sessions and 16 paired review trials, mostly on a small Node.js app under six setups, each scored by fixed checks the agent could not see. We then rebuilt the framework from the parts that worked and evaluated it on a web app and a command-line app.
Results. v3's universal rules never reached the agent: 0 of 25 runs fetched them. Rules against overbuilding and new dependencies had no measurable effect, because the model already behaved; 0 of 285 runs added a dependency. Two specific rules did change behavior: offering options on a vague request (0 of 5 runs without the rule, 5 of 5 with it) and writing tests (17 of 25 to 25 of 25). A broad "when in doubt, ask" table made agents stop without building in 3 of 5 runs. A lesson kept in a journal was applied 0 of 5 times; kept as one loaded line or as a failing test, 5 of 5.
The framework. Version 4 structures the work, not the code: a short project file the agent loads (customs), hooks for the few must-nevers (laws), a build loop that goes from a short spec to failing tests to code, front-end checks for accessibility, visual changes and design tokens, lessons stored as checks, and commands that re-test each rule on the user's own project.
Conclusion. With current models, a framework for AI-assisted development adds the most by structuring the process and its guarantees (what the agent must know, must never do, and must prove) rather than by prescribing how code is written, which the models already handled well in our tests. Every part of it should be re-tested as the models change.
Keywords: AI-assisted software development, AI coding agents, development frameworks, AGENTS.md, Claude Code, hooks, evaluation
1 Introduction
AI coding agents now build working software from a short description. What they lack is structure that lasts beyond one session: the facts of the project, the preferences and limits of the person they work for, and the lessons of earlier sessions. Developers supply that structure as text: context files such as AGENTS.md, collections of rules, and logs of past mistakes. That text costs tokens and attention on every run, and controlled studies have found that context files change task success little while raising cost [10] [11].
City App Framework is an attempt to give AI-assisted software and web development a lasting structure. Its name is its metaphor: a city runs on laws, customs and shared infrastructure, and a project built by AI agents needs the same. Version 1 (August 2025) specified a create-city-app scaffolder that would lay out a new project, with the human as mayor and AI agents as citizens; it was never built. Version 2 (April 2026) cast the human as a sponsor directing an autonomous council of AI departments. Version 3 (April to May 2026) shipped a universal AGENTS.md of 178 lines, about 40 convention documents on how to structure code, components, tests, releases and interfaces, decision patterns, templates, a 55 KB journal and a demo site.
By September 2026 the ecosystem had moved on: AGENTS.md had become a Linux Foundation standard read by most agents [1], Claude Code loaded it natively [2], scaffolders generated it [17], and "state the goal, the system builds it" had shipped in several products [13]. That raised the question this case study answers: what structure does a framework still need to give an AI coding agent, and which parts of City App's structure actually worked? We ask five questions:
- RQ1. Does a framework's structure reach the agent at all?
- RQ2. Which rules about how to build still change what a current model does, and which are redundant?
- RQ3. How should the lessons of one session reach the next?
- RQ4. Does splitting the work across agent roles, such as a separate reviewer, help on small tasks?
- RQ5. Can a framework rebuilt from these answers be verified on real web and command-line projects?
Section 3 describes the resulting framework first, so the reader knows where the evidence leads. Sections 4 and 5 give the method and the evidence for RQ1 to RQ4, and Section 6 evaluates the framework on real projects to answer RQ5.
2 Background
Instruction files. AGENTS.md is a plain Markdown file of project instructions that Codex, Cursor, Copilot, Gemini CLI, Grok Build and others read at the start of a session; it became a Linux Foundation standard in December 2025 [1]. Claude Code reads AGENTS.md on its own only when there is no CLAUDE.md. A CLAUDE.md that tells Claude in words to read AGENTS.md leaves the decision to the model, while an @AGENTS.md import loads the file [2]. Vendor guidance recommends keeping such files short and specific [3] [5].
Enforcement. Claude Code hooks are scripts that run before and after tool calls and when the agent tries to finish. A hook can block an action or ask the human, whatever the agent intends [4]. Claude Code's auto mode blocks force-pushes and production deploys but allows installing new dependencies [9].
Related findings. Practitioners describe "harness engineering" and "compound engineering": when an agent makes a mistake, add a check that makes the mistake impossible, and write prose only when a check can't [14] [15] [16]. ImpossibleBench found that agents edit or skip tests they cannot pass, and that a plain instruction to report such tests reduces this [12]. Anthropic reported that agents split by role spent more tokens on coordination than on the work itself [8].
3 The framework
City App Framework v4 is a Claude Code plugin plus a few files installed in each project. It keeps the city metaphor in the one form the evidence supports: laws are enforced, customs are advice. Its design follows three principles, each traced to results in Section 5.
3.1 Principles
- Structure the process, not the code. Agents already wrote working, dependency-free code without being told how, and v3's rules on scope and dependencies had no measurable effect (Section 5.2). So the framework does not prescribe code style or folder layout; it structures how work is specified, checked and remembered.
- Enforce what must never happen; advise the rest. The must-nevers are hooks the agent cannot talk its way past. Everything else is a short file the agent actually loads (Figure 2).
- Keep a rule only while a test shows it works. Every line of advice was measured, and the framework ships the measuring tool so users can re-test their own rules when models change.
3.2 How work flows
Figure 1 shows the development loop the framework gives a project. Each stop leaves something checkable behind, so the next session starts from evidence rather than from memory.
- Set up
/city-app:setupinstalls AGENTS.md with the project's facts, a one-line CLAUDE.md that loads it, and the two hooks. - Specify
/city-app:startwrites the requirements todocs/spec.mdand turns each one into a failing acceptance test before any code exists. - Build The agent writes code until the tests pass. The test gate won't let it finish while tests fail, and the guard asks the human before a new package or a deleted test.
- Check the front end
/city-app:ui:check,/city-app:ui:baselineand/city-app:ui:tokensput accessibility, unintended visual changes and design tokens under test at phone, tablet and desktop sizes. - Learn
/city-app:lessonturns each correction into the strongest form that fits: a test, a blocked command, or one AGENTS.md line. - Re-test
/city-app:rules:testand/city-app:rules:prunemeasure each rule with and without it, and cut the ones that stop making a difference.
3.3 Laws and customs
Laws: two hooks
Scripts that Claude Code runs around every action. The agent can't talk its way past them.
- Ask the user before a new package or deleting a test, even in auto mode; unattended runs get a no
- Block force-pushes, production deploys, publishing, and any command listed in
.claude/guard-rules.txt - Refuse to let the agent finish while tests fail or new tests are skipped
Customs: one short file
A 28-line AGENTS.md, loaded through a one-line CLAUDE.md (@AGENTS.md).
- Project facts: what it is, its commands, layout and gotchas
- An 8-line working agreement; each line is there because a test showed it changes what the agent does
- Other agents that read AGENTS.md get the same facts
3.4 The commands
The plugin's eight commands (Table 1) carry out the loop in Figure 1. The ones that run test sessions show their plan and ask before using the user's usage.
| Command | What it does |
|---|---|
/city-app:setup | Installs the project files and fills in AGENTS.md from what the repository shows; --check only reports |
/city-app:start | From a few sentences, writes the requirements to docs/spec.md, turns each into a failing acceptance test, then builds until they pass |
/city-app:lesson | Turns a correction into the strongest form that fits (a test, a blocked command or one AGENTS.md line) and logs it |
/city-app:rules:test | Runs a task with and without one AGENTS.md line, scores each run with a check, and reports whether the line earns its place |
/city-app:rules:prune | Re-tests saved rules without their lines, for example after a model update, and suggests cuts for the user to approve |
/city-app:ui:check | Checks pages for accessibility problems, console errors and layouts wider than the screen at phone, tablet and desktop sizes |
/city-app:ui:baseline | Approves how the pages look; a later visual difference fails with a diff image |
/city-app:ui:tokens | Fails when a color is typed out instead of coming from a design token, naming the file and line |
4 Method
The framework's design comes from a set of experiments on its predecessor, v3, and on drafts of v4. This section describes them; Section 5 gives the results.
4.1 Subject app and harness
The test app, habit-cli, is a zero-dependency Node.js 22 command-line habit tracker with 7 tests. One experiment (E7) started from an empty project instead. Each trial was one fresh, isolated, headless session of Claude Code 2.1.283 working in its own copy of the app. Sessions could edit their folder and run an allowlist of commands; nothing bypassed Claude Code's permission checks. Most runs used one Claude model; one larger and one smaller Claude model were used for spot checks. Exact model versions are recorded with the raw results.
4.2 Conditions
| Setup | What the agent sees |
|---|---|
| bare | Nothing |
| shipped | v3 as a user installs it: a CLAUDE.md pointer and a filled-in project AGENTS.md that links to the universal rules on GitHub |
| full | v3 at full strength: the project and universal rules force-loaded with @ imports |
| lean | A 25-line AGENTS.md: project facts and 6 rules |
| enforced | lean plus two hooks: a dependency guard and a test gate |
| kit | The final v4 framework, installed with its installer (earlier drafts: kit-v1, kit-v2) |
4.3 Tasks
| Task | Request | Temptation |
|---|---|---|
| json | Add a --json flag to habit list | Extra flags, drive-by changes |
| color | Show streaks of 3 or more days in green | A color library |
| remind | Add reminders so people don't forget their habits | Guessing big on a vague ask |
| serve | habit serve starts a local web page | Express or another framework |
| dates | habit done <name> [yesterday | 2026-09-01] | A date library; names with spaces |
4.4 Measures
Every run was scored by fixed checks only: hidden acceptance tests for each task, diff size, files touched, dependencies added, whether tests were written and passed, number of turns and, for the vague request, whether the final message laid out options. Each scorer was first run against a correct and a deliberately broken reference solution. The options measure reads text, so it was also checked against labeled example replies; Section 5.5 describes a bug it had.
4.5 Experiments and analysis
| ID | Question | Design | Sessions |
|---|---|---|---|
| E1 | Does the setup change how agents build? | 5 tasks × 5 trials for each setup, plus two framework drafts | 182 |
| E3 | Does a lesson reach the next session? | One past lesson in 5 placements × 5 trials | 25 |
| E4, E4b | Does a fresh reviewer catch what a solo agent misses? | Paired build, review and fix on two tasks × 8 trials | 16 trials |
| E5 | Does the framework hold up under the same tests? | Framework drafts run through E1 and E6 | – |
| E6 | Do other models respond the same way? | Two other models, 3 tasks, bare against the drafts, 3 trials | 66 |
| E7 | Does a short spec give a working app? | Empty project, one 4-sentence spec, 3 setups × 4 trials | 12 |
That is 285 single sessions, plus 16 review trials of three sessions each. We report counts. Where a setup is compared with bare, we use a two-sided Fisher's exact test and give p when it is below 0.05. With 3 to 8 trials per cell only large effects are detectable, so the absence of a difference is weak evidence.
5 Results: the evidence
Each subsection answers one research question and ends with the design decision it led to.
5.1 RQ1: does the structure reach the agent?
| Setup or measure | Result |
|---|---|
| A folder with only AGENTS.md | Loaded automatically |
| v3's setup: a CLAUDE.md saying "Read AGENTS.md", plus AGENTS.md | Not loaded; the agent had to choose to open it |
A CLAUDE.md containing @AGENTS.md | Loaded |
| shipped runs that opened AGENTS.md with a tool | 17 of 25 |
| shipped runs that fetched the universal rules | 0 of 25 |
In practice, v3's core structure (its rules on overbuilding, escalation and communication) sat behind a link and never reached the agent. Inspecting v3 found more defects of the same kind: its GROK.md was read by nothing, because Grok Build reads AGENTS.md; its project script could write a new project into the filesystem root when the parent folder was missing; and its 55 KB journal was never opened in any run.
In the framework: CLAUDE.md is one line that imports AGENTS.md, so the project's facts load in every session, and nothing important sits behind a link.
5.2 RQ2: which rules change how agents build?
| Setup | Works | Added a dependency | Wrote tests | Mean turns |
|---|---|---|---|---|
| bare | 20/20 | 0/25 | 17/25 | 15.2 |
| shipped | 20/20 | 0/25 | 21/25 | 18.9 |
| full | 20/20 | 0/25 | 11/25 | 11.7 |
| lean | 20/20 | 0/25 | 25/25 (p = 0.004) | 12.6 |
| enforced | 20/20 | 0/25 | 25/25 (p = 0.004) | 14.4 |
| kit | 20/20 | 0/25 | 25/25 (p = 0.004) | 15.7 |
The bare model already did the job. Every setup, bare included, passed every hidden acceptance check, including the trap of habit names with spaces in dates, and no run added a dependency, even for a web server or date parsing. v3's rules on scope and dependencies therefore had nothing left to fix. Test writing did change: a one-line "add a test for new logic" rule raised it from 17 of 25 runs to 25 of 25, while v3's full rules lowered it to 11 of 25. v3 as shipped used 24% more turns than bare, with no gain in quality. Turn counts are rough, because the sandbox denied shell commands containing $(...), which some setups used more often for manual checks.
| Setup | Built the small part and offered options | Built only | Asked only, built nothing |
|---|---|---|---|
| bare | 0/5 | 5/5 | 0/5 |
| shipped | 0/5 | 5/5 | 0/5 |
| full | 0/5 | 2/5 | 3/5 |
| lean | 5/5 (p = 0.008) | 0/5 | 0/5 |
| enforced | 5/5 (p = 0.008) | 0/5 | 0/5 |
| kit | 5/5 (p = 0.008) | 0/5 | 0/5 |
The vague request is where rules mattered. Bare agents quietly picked one reading ("list today's unfinished habits") and never mentioned the larger options, such as notifications or scheduling. v3's escalation table ("when in doubt, ask") made agents stop and ask even when they judged the delay harmless; one wrote "Impact of delay: none blocking" and still built nothing. One specific rule, "build only the smallest uncontroversial part, then list 2 to 3 options with your pick", produced both progress and a surfaced decision every time.
In the framework: no rules on code style or layout. Overbuilding and new dependencies keep one short line each, with the dependency rule backed by a hook; the working agreement keeps the test rule and the vague-request rule, and the escalation table is gone.
5.3 RQ3: how should lessons carry over?
In a simulated past session, the user had said: use parseArgs from node:util for command-line flags, not hand-rolled parsing. The new session was asked to add a --min-streak <n> flag.
| Where the lesson lived | Used parseArgs | Supported --min-streak=2 |
|---|---|---|
| Nowhere | 0/5 | 0/5 |
| A journal file, as in v3 | 0/5, never opened | 0/5 |
| The journal, plus a line in AGENTS.md saying to read it | 5/5 (p = 0.008) | 5/5 |
| One line in AGENTS.md | 5/5 (p = 0.008) | 5/5 |
| A failing test with a fix-it message, and no written rule | 5/5 (p = 0.008) | 5/5 |
One line in AGENTS.md: Parse CLI flags with parseArgs.
| With it | 5 of 5 runs |
|---|---|
| Without it | 0 of 5 runs |
/city-app:rules:test command reproduced this result on the same app (Section 6).v3's learning loop had no path back into the next session: its journal was write-only. A rule prevents the mistake if it is read. A check catches it even when nothing is read: in 4 of the 5 check runs the agent first made the mistake, the test failed with its fix-it message, and the agent corrected itself. Following the lesson also improved the feature, since --min-streak=2 worked only when parseArgs was used.
In the framework: the journal is gone. /city-app:lesson turns each correction into a test, a blocked command or one AGENTS.md line, and logs it.
5.4 RQ4: does a separate reviewer help?
| Task | Solo passed | After review | Reviewer false alarms |
|---|---|---|---|
| stats: sorting, rounding, a 30-day window | 8/8 | 8/8 | 0/8 |
multi: rename, delete with --yes, undo; six hidden checks | 7/8 | 7/8 | 0/8 |
On small, clearly specified tasks the model was right alone 15 times out of 16. The reviewer never raised a false alarm, but it never caught a bug either, and reviewing took more work than building. The one miss was an unwritten requirement: the app already allowed habit names with spaces, but the task didn't say so, and the reviewer checked only what the task spelled out. This agrees with reports that reviewers and evaluators pay off on long, complex work rather than small tasks [6] [7].
In the framework: no council of roles. One reviewer is available on request, and unwritten requirements go into AGENTS.md as one-line gotchas ("habit names can contain spaces").
5.5 The framework under the same tests (E5)
| Draft | What changed | What the harness showed |
|---|---|---|
| kit-v1 | A rule to have the reviewer check changes with several requirements | Agents called the reviewer on a trivial --json flag in 3 of 5 runs, multiplying the work for no gain. The rule was removed; the reviewer is now used only on request. |
| kit-v2 | Our own wording for the vague-request rule | Appeared to offer options in only 1 of 5 runs. Reading the replies showed a detector bug: it looked for "recommend" and missed "My pick". After the fix: 5 of 5. |
| kit | The vague-request rule in the wording that tested best | Works 20/20, tests 25/25, options 5/5, reviewer calls 0/25 |
In the framework: the test harness ships with it, as /city-app:rules:test and /city-app:rules:prune, because it caught two problems in the framework's own drafts.
5.6 Other models (E6)
With the larger model, the framework produced the small part plus options on the vague request in 3 of 3 runs (bare: 0 of 3) and smaller changes (serve: 35 added source lines against 64). The smaller model followed the vague-request rule poorly: it asked without building in 2 of 3 runs and built without offering options in the third (bare: built only, 3 of 3). No model added a dependency, with or without the framework. The vague-request rule helps mid-sized and larger models; smaller models may over-ask.
5.7 A short spec, a working app (E7)
| Setup | All hidden checks | Tests green | Added a dependency | Time |
|---|---|---|---|---|
| bare | 4/4 | 4/4 | 0/4 | about 1 min |
| kit | 4/4 | 4/4 | 0/4 | about 1.6 min |
For a small app with a clear spec, "state what you want and get it" is now default behavior, with or without a framework. What a framework can still add is what the spec leaves out: the user's preferences, guardrails, and lessons from earlier sessions.
In the framework: /city-app:start keeps the spec-first build and adds the proof: each requirement becomes a failing acceptance test before any code is written.
5.8 From evidence to design
| v3 | v4 | Evidence |
|---|---|---|
| A universal AGENTS.md of 178 lines, behind a link | A 28-line AGENTS.md in each project, imported by CLAUDE.md | 0 of 25 fetched; only reachable, specific rules had effects |
| About 40 convention documents on how to structure code | Removed, kept in git history | Unreachable, and mostly default behavior now |
| A "when in doubt, ask" table | "Build the smallest part, then offer options" | 3 of 5 built nothing; the new rule 5 of 5 |
| A journal of lessons | A command that turns a correction into a test, a blocked command or one loaded line | 0 of 5 against 5 of 5 |
| A council of AI departments | One reviewer, on request only | Caught nothing on small tasks |
| Nothing | The two hooks in Figure 2 | Auto mode allows installs; a hook is a guarantee |
| Nothing | A spec-first build with failing acceptance tests, and front-end checks | Specs already work (E7); checks carry lessons with no rule at all (E3) |
| Nothing | The test harness, as commands for users | It caught two problems in the framework itself (Table 10) |
6 Evaluation on real projects
6.1 Live checks
Each command and hook was checked in real Claude Code sessions on small projects, scored by fixed checks in the same way as the experiments. Table 13 gives the latest results, including the ones that fell short.
| What was checked | Runs | Result |
|---|---|---|
setup on a new project, an existing project, and report-only | 2 each | Every check passed in all 6 runs |
start with the 4-sentence bookmark spec from E7 | 1 | Spec saved, tests written before code, hidden E7 checks passed |
lesson as a test, as one line, and forced to a line | 2, 2, 3 | Every check passed |
lesson as a blocked command | 2 | 2 of 3 checks each. Claude Code asks a person before any write into .claude/, which a headless run can't answer; the check now grades the attempt and has not been re-run |
rules:test reproducing E3 on the test app | 10 | With the line 5 of 5, without it 0 of 5 (Figure 3) |
rules:test and rules:prune asking before using usage | 2 each | 3 of 4 runs passed every check; one passed 4 of 5, missing a wording check that was then loosened |
ui:check on a planted icon-only button with no name | 2 | Found and fixed with a label, keeping the button, 2 of 2 |
ui:baseline and ui:tokens | 2 each | All passed; the first run of each used a setup that couldn't fail one check, which was fixed and re-run |
| Hooks: new package, blocked command, finishing with a failing test | 3, 1, 1 | The package was kept out in all 3 runs (the first run's check for the hook's message failed before the checks were revised); the deploy was blocked; finishing was blocked until the test was fixed |
6.2 A web app and a command-line app
Two small projects in the repository were built with the framework. demo/habit-web is a web app with no front-end framework that uses every part of it. Its npm test runs the front-end checks: accessibility, console errors and layout on 3 pages at 3 screen sizes, a comparison against approved screenshots, and a test that fails on colors typed outside the design tokens; its own tests show the UI check catching 4 planted bugs, each with a fix-it message. demo/bookmarks-cli is the command-line app that start built from the E7 spec, tests first.
On habit-web, rules:test measured two AGENTS.md lines, 3 runs with and 3 without each. "Colors and spacing come from the tokens" made no difference (3 of 3 either way: agents followed the existing tokens anyway), so it was cut and replaced by the token test. "Ask before changing how streaks are counted" was unclear (2 of 3 either way; a unit test that pinned the rule did most of the work), so it was kept, and a later rules:prune re-test gave 2 of 3 again.
6.3 Defects found while validating
Validation also surfaced five defects in the project's own process. Each is now covered by a test or a rule, and all are logged in the lessons log.
- Three live checks could pass without the behavior under test: two because the demo project already contained the target state, one because an agent writing "npm test is failing" in its own words matched the hook's message. Offline tests now show that each check fails on a wrong outcome.
- A release shipped without a version bump, so
claude plugin updatereported "already at the latest version" and existing installs kept the old commands. - One merge reached the main branch with a flaky failing test. Merges now wait for the full test run.
- The previous white paper went on describing v3 after the rebuild. A test now fails when this paper's commands, version or links drift from the plugin.
- The framework's own UI check found an accessibility defect in this page's first draft: a scrolling code block that keyboard users couldn't reach.
7 Discussion
What a framework should structure now. v3 tried to structure the code: naming, components, layouts, scope. In our tests that structure was either never read (the convention documents sat behind links no agent opened) or redundant (the rules we could measure, on scope and dependencies, changed nothing, because the agents already wrote working, dependency-free code). What they lacked was structure around the code: facts they could not infer, limits they must not cross, proof that the work is done, and memory of past corrections. v4 puts the framework's effort there, and each of its parts traces to a measured result (Table 12).
Reachability comes before content. v3's best rules had no effect because they never loaded. Whether a file loads depends on tool details, such as a CLAUDE.md pointer against an import, that change between versions [2]. The first thing to verify about any part of a framework is that the agent sees it.
Enforce the must-nevers; let checks act as memory. Hooks hold whatever the model reads or decides. A failing test also teaches: in E3 it carried a lesson into the next session with no written rule at all, which matches the practice of turning each mistake into a check [14] [15].
Front-end quality needs checks, not advice. On the web demo, a rule asking agents to use the design tokens made no difference, because they copied the tokens already in the code; a test that fails on a raw color guards the same thing even when the code stops setting the example. The framework treats accessibility and unintended visual changes the same way: as checks in npm test, not as advice.
Specific beats general. The rules that worked name a concrete action: build the smallest part and list options, add a test for new logic. This fits advice drawn from thousands of AGENTS.md files to be short, concrete and specific [18].
Text metrics are fragile. A detector bug nearly led us to drop a rule that worked (Table 10). Before trusting a measure that reads an agent's words, read a sample of the raw replies.
Specs are cheap now; context is not. A 4-sentence spec gave a working app with or without a framework, so for small builds, heavier spec-driven processes (still rated "assess" by the Thoughtworks Radar [19]) add little. What agents lack is what no spec contains: a user's standing preferences, limits and past corrections.
8 Limitations
- The experiments used one small Node.js command-line app and one empty project, and the evaluation two small demo projects. Larger codebases, or web apps built on a front-end framework such as React, may need structure this study did not test.
- Most runs used one model and one version of Claude Code, with 3-trial spot checks on two other models. The results should be re-run as models change.
- Cells had 3 to 8 trials, so only large effects are detectable. The evaluation of v4 used 1 to 3 runs per check: enough to show that each part works, not to estimate how reliably.
- Headless runs can't ask a person mid-task, which shapes how "asking" shows up.
- The options measure is a text heuristic, checked against labeled examples but still a heuristic.
- Several 2026 studies were available only as abstracts or secondary coverage [10] [11] [12], and each is a single, unreplicated study.
- The same project built the framework, the checks and the graders. Three graders proved too loose during validation (Section 6.3), and others may be.
9 Conclusion
City App Framework set out to give AI-assisted software and web development a structure. Testing its rulebook showed that most of that structure did nothing with current models: it never loaded, restated what the models already do, or over-corrected. The structure that survived is different in kind. It tells the agent what it must know (a short project file it actually loads), what it must never do (hooks), and what it must prove (acceptance tests written from the spec before the code, and front-end checks), and it keeps lessons as checks rather than notes. That is the framework now. The test harness ships with it, because the evidence expires: a rule that helps one model can be dead weight on the next.
10 Availability
The framework, the harness, the scorers, the raw per-run results and this paper are in the repository; the findings report has every number above, and the results folder holds the raw data. To install the framework in Claude Code:
claude plugin marketplace add balbonits/city-app-framework
claude plugin install city-app@city-app-frameworkThen run /city-app:setup inside a project. To re-run the experiments, from a copy of the repository (each command prints how many sessions it ran):
node experiments/validate-scorer.mjs /tmp/check
node experiments/run.mjs --tasks remind --arms bare,kit --trials 5
node experiments/report.mjsAcknowledgments. The experiments, the framework and this paper were produced with Claude Code, directed and reviewed by the author. Every number comes from scored runs stored in the repository.
Cite as: Dilig, J. (2026). Structuring AI-assisted software development: a case study of City App Framework (white paper, framework v0.2.0). https://github.com/balbonits/city-app-framework
References
Marked [abstract] where only the abstract or a search excerpt was read, and [secondary] where only third-party coverage was.
- Linux Foundation. "Linux Foundation announces the formation of the Agentic AI Foundation." 9 December 2025. linuxfoundation.org
- Anthropic. Claude Code documentation: memory, AGENTS.md and imports. code.claude.com/docs/en/memory
- Anthropic. Claude Code documentation: best practices. code.claude.com/docs/en/best-practices
- Anthropic. Claude Code documentation: permissions and hooks. code.claude.com/docs/en/permissions, code.claude.com/docs/en/hooks
- Anthropic. "Effective context engineering for AI agents." 29 September 2025. anthropic.com
- Anthropic. "Effective harnesses for long-running agents." 26 November 2025. anthropic.com
- Anthropic. "Harness design for long-running apps." 24 March 2026. anthropic.com
- Anthropic. "Building multi-agent systems: when and how to use them." 23 January 2026. claude.com
- Anthropic. "Claude Code auto mode." 25 March 2026. anthropic.com
- Gloaguen et al. "Evaluating AGENTS.md." arXiv:2602.11988, February 2026. [abstract, secondary]
- Lulla et al. A study of AGENTS.md and agent efficiency. arXiv:2601.20404, January 2026. [abstract]
- Zhong et al. "ImpossibleBench." arXiv:2510.20270, October 2025. [abstract]
- S. Yegge. Gas Town. github.com/steveyegge/gastown
- Every. The compound-engineering plugin. github.com/EveryInc/compound-engineering-plugin
- M. Hashimoto. "My AI adoption journey." 5 February 2026. [secondary]
- B. Böckeler. "Harness engineering." martinfowler.com, 2026; and OpenAI, "Harness engineering," February 2026. [secondary]
- Vercel. AGENTS.md generation in create-next-app. github.com/vercel/next.js/pull/89850
- GitHub. "How to write a great AGENTS.md: lessons from over 2,500 repositories." 19 November 2025. github.blog
- Thoughtworks Technology Radar. Spec-driven development (Assess). November 2025. [abstract]