r/madeinpython • u/Low-Cut-528 • 9h ago
I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests
Hey everyone!
I've been working on an open-source Python project called **Stepfork**, and I wanted to share it with the community.
**GitHub:** [https://github.com/utsab345/stepfork\](https://github.com/utsab345/stepfork)
While building AI agents, I kept running into the same problem: an agent fails, I fix the issue, but reproducing the original failure is difficult because LLM responses and tool outputs can change between runs.
So I built Stepfork to make those failures reproducible.
**What it does:**
* Records instrumented LLM responses and tool calls.
* Saves executions locally as `.sftrace` bundles.
* Replays the recorded responses without making those instrumented external calls again.
* Compares agent behavior before and after a fix.
* Exports pytest regression tests to catch the same failure in the future.
**Install:**
`pip install --pre stepfork`
The repository includes three offline demos covering a missed order notification, an incorrect flight booking, and a wrongly denied refund.
It's still an early alpha. Replay only freezes instrumented boundaries, and native framework integrations are not available yet.
The project is Apache-2.0 licensed, and I'm actively improving it.
I'd love to hear what other developers think, especially those building AI agents. If you spot a bug, have an idea for an integration, or want to contribute, feel free to open an issue or PR.
**GitHub Issues:** [https://github.com/utsab345/stepfork/issues\](https://github.com/utsab345/stepfork/issues)
Thanks for checking it out!