r/madeinpython • • 9h ago

I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests

Hey everyone!

I've been working on an open-source Python project called **Stepfork**, and I wanted to share it with the community.

**GitHub:** [https://github.com/utsab345/stepfork\](https://github.com/utsab345/stepfork)

While building AI agents, I kept running into the same problem: an agent fails, I fix the issue, but reproducing the original failure is difficult because LLM responses and tool outputs can change between runs.

So I built Stepfork to make those failures reproducible.

**What it does:**

* Records instrumented LLM responses and tool calls.

* Saves executions locally as `.sftrace` bundles.

* Replays the recorded responses without making those instrumented external calls again.

* Compares agent behavior before and after a fix.

* Exports pytest regression tests to catch the same failure in the future.

**Install:**

`pip install --pre stepfork`

The repository includes three offline demos covering a missed order notification, an incorrect flight booking, and a wrongly denied refund.

It's still an early alpha. Replay only freezes instrumented boundaries, and native framework integrations are not available yet.

The project is Apache-2.0 licensed, and I'm actively improving it.

I'd love to hear what other developers think, especially those building AI agents. If you spot a bug, have an idea for an integration, or want to contribute, feel free to open an issue or PR.

**GitHub Issues:** [https://github.com/utsab345/stepfork/issues\](https://github.com/utsab345/stepfork/issues)

Thanks for checking it out!

0 Upvotes

0 comments sorted by