Artificial Intelligence
AI Agents
·
This walkthrough uses a real Apollo Kotlin bug to compare how two agent setups investigate a problem, change code, and demonstrate that the change is safe. You can run the task yourself from the exact pre-fix revision, or replay the saved solutions included with this article.
The result is close. Both candidates fix the cache-read defect with essentially the same algorithm. I prefer the Coral-assisted submission by a modest margin because it tests more realistic cache states and checks more of the surrounding cache behavior. Plain Codex contributes an especially good regression for fragment order. For the production implementation itself, I prefer the smaller change that Apollo’s maintainers merged.
What Coral adds
coral-code is a JetBrains IDE plugin that lets a coding agent work with a graph of agents investigating the codebase. The outer agent can request analysis, read findings as they arrive, and follow up while continuing its own source inspection and tests. It is available on the JetBrains Marketplace.
For broader context, Coral’s public baseline study reports these results on 124 questions across 11 repositories, using its official-rubric evaluation:
Configuration | Perfect answers | Rate |
|---|---|---|
Coral | 95 / 124 | 76.6% |
Opus 5.5 xhigh, strongest single-model baseline | 64 / 124 | 51.6% |
DeepSeek V4.1 Flash alone | 63 / 124 | 50.8% |
That is a 25.0 percentage-point lead over the strongest single-model baseline. These are Coral-published codebase Q&A results, dated September 29, 2026; they are context for this case study, not a patch-success rate or a score for the Apollo task. See the public results and methodology and its linked reproduction repository.

Here the question is qualitative: which solution would you rather maintain, and what evidence makes you trust it?
1. Start immediately before the fix
The issue is Apollo Kotlin #6901, “Fetch failures with @include directive”. The reporter used response-based models and overlapping fragments. Parsing the network response worked; reading the data back from the normalized cache failed. That cache-only clarification defines what the regression must exercise.
Use this exact baseline:
It is the sole parent of merged fix 0e70dff22fac954403d47c62f054f24be00dcc64, from PR #6922. Checking out today’s default branch would give the agents the already-fixed reader.

You need Git, a JDK 17 or newer, and network access for Gradle dependencies. The commands use a macOS/Linux shell; Windows users can use WSL. The optional replay script also needs Python 3.9 or newer. JDK 21 is a reasonable common choice. Use the repository’s Gradle wrapper; this revision pins Gradle 9.4.0. The walkthrough runs JVM tests, so it does not require Xcode or a simulator. Use a compatible JetBrains IDE for the Coral run. Apollo’s pinned contributor guide describes the composite build.
Create two ordinary clones in a new comparison directory:
Both hashes should match BASE; both status commands should initially be empty. Run the agents sequentially if that is more comfortable for your machine. Simultaneous execution is unnecessary.
Keep this article, its answer snapshots, and the other candidate’s outputs outside each agent’s project context until the runs finish.
2. Configure the two runs
For Codex alone, open apollo-plain in your usual Codex client and start a new chat. Disable Coral tools for that chat. Do not give it an earlier conversation or the solution sections below.
For Codex with Coral, install Coral Code from the JetBrains Marketplace and restart the IDE if prompted. Open apollo-coral, finish project indexing/import, and open View → Tool Windows → coral-code Chat. Select Codex as the outer agent, connect its model account, configure the graph provider/account, and enable graph assistance. Use the plugin’s installed Coral skill/instructions for this project. If the release offers a graph mode, record the one you choose; High certainty is a reasonable choice for this investigation. See the installation and configuration documentation.
Choose the same outer model and reasoning effort for both runs. Record the model identifier, reasoning setting, Codex version, Coral plugin version and graph configuration, IDE version, JDK, and operating system. These settings change over time; a reproducible comparison should name the versions actually used. The bundle includes a run-notes template.

Give both agents the shared prompt below, followed by their short run-specific instruction. This compares the two working setups, including their available tools and context. It is a fresh trial protocol, not a claim that the historical chats had identical starting conditions.
Shared task prompt
Append for the plain run:
Append for the Coral run:
Save the final diff, all new source/test files, test commands and reports, and the chat transcript. git diff alone omits untracked tests. For Coral, also retain evidence that a question was submitted and findings were read during the work. It is fine to work from provisional findings; an outer turn need not wait for terminal graph publication to make a supported decision.
3. Replay the preserved solutions
Fresh agent runs can produce different valid patches. The companion bundle therefore includes the complete saved candidate files, full patches including new tests, the exact upstream changed files, and a replay script. Keep the bundle’s directory structure intact when downloading or sharing it.
From the directory containing this README:
Each command requires a new destination directory. It clones upstream, checks the pinned commits’ parent relationship, verifies snapshot checksums, and applies the candidate’s complete patch. It then runs the same regression tests with three reader implementations:
The original baseline reader, expecting the specific regression failures.
That candidate’s reader, expecting every regression to pass.
The merged upstream reader, keeping the candidate’s tests, expecting every regression to pass.
The third step adds a cross-check beyond the historical agent runs. All six replay phases were executed while preparing this bundle: the baseline failures reproduced, each candidate passed its own tests, and the upstream reader passed both suites—4/4 plain and 5/5 Coral. See the fresh replay results, recorded separately from the historical executions below.
The script retains each phase’s Gradle log and JUnit XML, and writes results.json. Missing tests, skipped cases, an unexpected failing method, or a build failure before testing do not count as a reproduced red/green result. A stopped run leaves its new clone available for inspection.
For a lightweight preparation check without compiling, use a separate new destination:
The replay uses this command shape from its newly created clone:
For Coral’s saved tests, the class is test.OverlappingFragmentsTest. The included init script disables the configured remote build cache and build-scan publication. At this revision, configuration on demand also avoids an unrelated Apple-target configuration failure encountered during the original JVM-only run. These settings make the local build easier to reproduce; they are not part of either fix.
For a broader common check, run these tasks on each completed candidate using the same JVM-only environment and Gradle flags above, but omit --tests:
For targeted integration checks, use -p tests :integration-tests:jvmTest with --tests test.NormalizerTest --tests test.StoreTest --tests test.OtherCacheTest --tests test.CacheResolverTest. The JVM-only test versus multiplatform jvmTest distinction matters: not every module exposes both tasks.
4. Understand the failure before comparing the patches
The query selects details twice. With includeExtras=true, both occurrences are active, and their child fields must survive:
The baseline reader groups selections by response name and directive condition:
That keeps unconditional details separate from conditional details, even when the condition is true. Both reads reach the same response path. Later, the reader stores each reconstructed object’s map with a plain assignment:
The second partial map overwrites the first. In the supplied fragment order, required title disappears and the read throws. Reverse the fragments and nullable tags can disappear silently. The normalized record can contain all the right fields while reconstruction still loses data.

This is why both candidates check parsing and the stored record before asserting the cache-read result. A parse-only test would miss the defect. The separate conditional-nullability/code-generation concern tracked in #6923 is outside this fix’s scope.
5. What Codex alone produced
Plain Codex followed a direct source-to-test investigation. It identified the duplicate-path overwrite, related read behavior to the existing normalizer, and reproduced the failure before editing production code.
Its patch filters skipped fields during recursive collection:
Then it merges the active selections by response name and clears their already-evaluated conditions:
Its strongest test choice is reversing fragment order. The four cases are included/excluded selections, each in the original and reversed order. All use explicit ID-based cache keys. Checking complete model equality catches both the exception and silent nullable-field loss:
That is a strong mechanistic test: it demonstrates an overwrite rather than merely reproducing one error message. See the complete plain patch and test class.
6. What Codex with Coral produced
The Coral-assisted candidate reaches the same effective production algorithm. It leaves recursive field collection unchanged and places filtering next to merging in one helper:
Both candidates use their helper for normalized records and the nested-map path, preserve fragment type checks, and keep aliases distinct by grouping on response name. Neither changes code generation or adds a dependency. Both resemble the writer’s existing practice of filtering active selections and merging their children in Normalizer.kt. I slightly prefer Coral’s localized transformation for readability, but the algorithms are effectively tied for this defect.
The more meaningful difference is validation. Coral’s five cases cover true/false conditions with both explicit ID keys and default path-based keys, plus a false-condition read after a true-condition write:
That last case asks something a fresh false-condition round trip cannot: does the read still respect its selection when extras already exist in the cache? The path-key cases also avoid validating only applications with a custom ID generator. See the complete Coral patch and test class.
Was the graph actually used?
Yes. The retained session shows a graph request, findings delivered while the outer agent was working, and a substantive follow-up responding to them:
Point in the retained turn | Observed interaction |
|---|---|
04:00:59 UTC | The outer agent submits a graph question. |
04:06:41 | Reader findings reach the outer agent. |
Shortly afterward | A provisional synthesis raises a possible path-collision concern. |
04:07:31 | The outer agent replies with source reasoning and test evidence, explaining that the apparent |
04:12:54 onward | The outer agent reports its tested result; terminal graph publication follows seconds later, and a subsequent turn reads and reconciles it. |
The graph participated in the investigation and review. The outer agent also challenged an incorrect hypothesis rather than accepting it automatically. Working with provisional findings is part of this workflow; terminal publication arriving after the first final answer is not itself a quality problem.

One historical detail matters: the retained Coral turn started with a candidate and three tests already present, then verified the baseline and added path-key coverage. The plain trace starts from a clean checkout. We can compare the resulting code and demonstrated validation, but those starting states make elapsed-time comparisons unhelpful. The fresh-run instructions above give both agents a clean baseline.
7. Results and qualitative judgment
These are historical executed results, corroborated by saved logs, test reports, and tool output. They are not newly executed results from preparing this article:
Candidate regression | Original reader | Candidate reader |
|---|---|---|
Plain Codex | 2 passed, 2 failed out of 4 | 4 / 4 passed |
Codex with Coral | 3 passed, 2 failed out of 5 | 5 / 5 passed |
Plain’s two failures expose missing title and silent tags loss under reversed order. Coral’s failures expose missing title with ID and path-based keys. In both cases, parsing and raw-record assertions pass before the cache read fails.

The surrounding suites also passed:
Run | Recorded broader validation |
|---|---|
Plain | Response-based models 13/13; operation-based models 9/9; cache variables/arguments 7/7; normalized-cache API 32/32 |
Coral | Response-based models 14/14; cache variables/arguments 7/7; normalization 6/6; include/skip 12/12; selected cache integration tests 27/27 |
Focused regressions are already included in the response-based totals; do not add them twice. The differing suite choices are useful engineering evidence, not a matched numerical leaderboard.
Factor | Judgment |
|---|---|
Correctness on #6901 | Tie: both fix the demonstrated cache-read loss. |
Fit with the codebase | Tie overall: both reuse an existing selection-merging pattern and existing test infrastructure. |
Order dependence and silent loss | Plain wins: reversing fragments tests a particularly sharp consequence of the mechanism. |
Cache states and key strategies | Coral wins: default keys and populated-cache reads represent meaningful additional usage conditions. |
Side-effect checks | Coral has more directly relevant integration coverage; plain adds useful breadth across generated-model styles. |
Maintainability | Slight Coral advantage in localized filtering/merging and explicit test-store cleanup. |
Investigation workflow | Plain is more direct. Coral adds an active review dialogue, with some tool-recovery and coordination overhead. |
My preference is the Coral-assisted package, by a modest margin. I weight its cache-state coverage and surrounding integration checks more heavily than the small style differences. That preference does not depend on tracing every successful decision to an individual graph answer. It is a judgment about the submitted code, demonstrated behavior, and observed review process.

Plain’s reverse-order case is too useful to discard. I would add it to the Coral package before calling the regression coverage complete.
8. Compare the preferred candidate with the merged PR
Apollo’s merged solution fixes the same defect at a different point. The candidates merge active selections before traversal. Upstream merges the partial object data when accumulating the result:
The full upstream implementation explains why overlapping keys should contain the same value and why shallow merging is sufficient: the reader works with normalized records; nested maps here represent scalar values expected to agree.

I prefer this production change for three reasons:
It changes less observable behavior. The existing field-collection and resolver-call pattern stays intact. Both candidates combine selections, clear directive metadata, and reduce duplicate resolver calls. No failing custom resolver was demonstrated, but these are additional behaviors to assess in a library with extension points.
It establishes a local invariant. Repeated partial reads at one response path preserve the fields already collected. Correct accumulation remains true even if another valid route later produces repeated reads. The candidates retain the plain assignment and rely on earlier collection preventing duplication.
It documents the boundary. The shallow-merge explanation tells future maintainers what is assumed and why a deep merge is unnecessary.
Upstream’s regression is economical. It adds a TwoFieldsQuery(true) to the existing cache-directive fixtures, writes a model containing an unconditional user ID and conditional name, reads it back, and checks the name:
The complete maintainer test exercises the right mechanism with less fixture machinery. Coral’s tests stay closer to the reported response-based setup and cover more conditions.
To run the exact merged test in another fresh clone:
The strongest combined submission would use upstream’s implementation, Coral’s cache-state/key coverage, and plain Codex’s reverse-order regression. The fresh replay confirms that upstream’s reader passes each candidate suite independently. Consolidating those fixtures into one submission remains a maintenance choice, rather than a patch produced by either original agent run.

Record your own result
When repeating the exercise, keep the model/setup details, complete patches, before/after test reports, and short answers to these questions:
Did the test reproduce cache-read loss after parsing and normalization succeeded?
Does the fix respect directives, aliases, existing abstractions, and extension points?
Which plausible side effects were checked, and which remain uncertain?
What makes the code easy to maintain when another overlapping-selection case appears?
For Coral, were graph findings requested and read during the work, and how did the outer agent respond?
Rank the engineering artifacts first, then compare the preferred candidate with the pinned upstream fix. A useful case study can explain why one package is better without requiring a one-to-one attribution for every line of code.
The evidence notes describe the preserved sources and what was checked while assembling this bundle. Apollo-derived source is accompanied by its MIT license.


