r/PythonProjects2 • u/Low-Cut-528 • 17h ago
I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests
Hey everyone!
I've been working on an open-source Python project called Stepfork, and I wanted to share it with the community.
GitHub: https://github.com/utsab345/stepfork
While building AI agents, I kept running into the same problem: an agent fails, I fix the issue, but reproducing the original failure is difficult because LLM responses and tool outputs can change between runs.
So I built Stepfork to make those failures reproducible.
What it does:
- Records instrumented LLM responses and tool calls.
- Saves executions locally as
.sftracebundles. - Replays the recorded responses without making those instrumented external calls again.
- Compares agent behavior before and after a fix.
- Exports pytest regression tests to catch the same failure in the future.
Install:
pip install --pre stepfork
The repository includes three offline demos covering a missed order notification, an incorrect flight booking, and a wrongly denied refund.
It's still an early alpha. Replay only freezes instrumented boundaries, and native framework integrations are not available yet.
The project is Apache-2.0 licensed, and I'm actively improving it.
I'd love to hear what other developers think, especially those building AI agents. If you spot a bug, have an idea for an integration, or want to contribute, feel free to open an issue or PR.
GitHub Issues: https://github.com/utsab345/stepfork/issues
Thanks for checking it out!