The Tellenze blog

Leave the Failed Experiment in the Handoff

The attempt that didn't work can save someone else time. Preserve what it showed, where its limits lie and what would make it worth trying again.

· The Tellenze team

A brass magnifier enlarges the crack in a teal ceramic test piece, with a matching fragment beside it in an ivory-lined open wooden drawer.

The handoff note is nearly finished. The slow planning screen has been improved, the change is ready for review, and the engineer adds one last sentence:

We tried caching. It didn't help.

In this fictional investigation, that sentence refers to something quite specific. The engineer tested a planning screen with a large set of sample work items. Caching a calculation made little difference to the wait. The browser trace showed that laying out all those rows took most of the time.

The team chose to render fewer rows at once. The caching experiment was set aside.

A few months later, another engineer picks up a new performance complaint. The application now does more work to calculate what each person can see. They read the old note.

Should they skip caching? Try it again? Find the original engineer?

The experiment was useful. The sentence has made it difficult to use.

A failure can disappear in two directions

The first is straightforward: nobody records it. The handoff contains the final implementation, perhaps a test result, and a clean account of the successful approach. The discarded attempts stay in someone's memory.

The next person follows the same plausible route and meets the same obstacle. They may even feel pleased with the idea until someone says, “We already tried that.”

The other loss happens when a specific result becomes a general rule. “This calculation wasn't the bottleneck in that version” turns into “Caching doesn't help this screen.” People remember the prohibition and forget the experiment.

Weng Jialin's account of a long-running agent optimization task offers a concrete example of keeping unsuccessful work useful. The experiment registry retained candidates, hypotheses, results, validation and failure reasons so later agent sessions could recover the research state. The account describes one performance challenge, rather than proving that a particular note format works everywhere.

What interests me is the place given to the losing candidates. They remain part of the investigation.

For the planning screen, the failed caching trial deserves a few lines beside the successful change.

Keep the experiment small enough to understand

A useful handoff might say:

Aim: Make the large planning screen respond promptly.

Tried: Cache the row-summary calculation in the recorded build, using the linked large-project fixture.

Observed: Repeated runs showed little change in the overall wait. The attached browser traces put most of the delay in row layout.

Decision: Set aside this cache and reduce the number of rows rendered at once. The implementation and checks are linked; the change still awaits review.

Limit: This does not test caching of permission calculations, other project sizes or later versions. The patch and reproduction steps are retained.

Next: The receiving engineer can reproduce the current complaint and compare where its time goes before choosing an approach.

This is still an illustrative record, with no real benchmark result behind it. In actual work, identify the build, fixture, environment and measurements precisely enough for someone to find them.

The receiving engineer can start with this short account and open the relevant traces or patch if the new question needs them. Routine edits can stay in the ordinary development history.

Describe the state of the discarded patch too. It may have incomplete correctness checks or belong to an old interface. Say so. The useful thing being handed over is an experiment, including what it failed to establish.

The same distinction applies outside performance work. A prototype that confused participants may reveal a problem with one label, one audience or one task. “Users don't want this feature” would be a much larger conclusion.

Let the result have a past tense

Now return to the later complaint.

The receiving engineer compares the new trace with the old one. In our fictional example, permission calculations have become a substantial part of the delay. The original trial never tested those calculations.

This time, caching is worth investigating alongside the earlier layout improvement.

A team can respect both findings without deciding that the first engineer was wrong.

Changing tools can produce a similar moment. In Anthropic's April 2026 engineering report on Managed Agents, the authors describe a context-reset workaround for one model's tendency to finish tasks prematurely. With another model, they found that behavior absent and the workaround unnecessary. That is a vendor's observation in its own harness, but it makes the practical question clear: does the condition that justified the old response still exist?

For a handoff, “revisit when…” is often more useful than “never do…”

The trigger should be concrete. A new model, a different workload, a changed interface or a corrected test might matter. Mere enthusiasm for an old idea is less persuasive.

Keep the earlier finding available when you add the new one. Rewriting the record to say “caching helps” would repeat the original mistake with the opposite conclusion.

Try the handoff from the other side

There is a limit to how much a careful record can do.

The LongMemEval-V2 preprint distinguishes recalling facts from recognizing environment-specific problems and assumptions that do not fit the current environment. Its evaluation uses collected web-agent histories and question answering; it doesn't measure whether your engineering team will complete a handoff faster. The distinction is nevertheless useful here: retrieving an old result and knowing how to apply it are different jobs.

Human handoffs have their own difficulties. Jian Zhao and colleagues' study of asynchronous investigative handoffs explored ways to make earlier investigations understandable. The authors also observed that inherited misunderstandings could mislead later investigators. Their study concerns document analysis, with limits on its participants and experimental setting. A more legible history still needs a thoughtful reader.

Before finishing a real handoff, ask the person receiving it to pick up one important failed attempt. Can they explain what was tried, find its evidence and say whether the current situation falls within its limits?

If they cannot, that gives you a specific gap to repair. Perhaps the trace is missing. Perhaps the comparison used a different dataset. Perhaps the note states a conclusion more confidently than the result allows.

In Tellenze, the work-item guide describes keeping decisions and progress in Conversation, with supporting evidence and linked outcome documents. Put the short account where the work already lives and link the material needed to inspect it. Give the receiving person a test fixture and reproduction steps they can actually use.

For our receiving engineer, the old experiment now provides a starting point. They can reproduce the new complaint, inspect the changed calculations and design a comparison around them.

Their next trial asks: “Would caching this calculation reduce the delay we can see now?”

Further reading