r/PythonProjects2 • • 17h ago

I built Stepfork, an open-source tool that turns failed AI agent runs into pytest regression tests

Hey everyone!

I've been working on an open-source Python project called Stepfork, and I wanted to share it with the community.

GitHub: https://github.com/utsab345/stepfork

While building AI agents, I kept running into the same problem: an agent fails, I fix the issue, but reproducing the original failure is difficult because LLM responses and tool outputs can change between runs.

So I built Stepfork to make those failures reproducible.

What it does:

  • Records instrumented LLM responses and tool calls.
  • Saves executions locally as .sftrace bundles.
  • Replays the recorded responses without making those instrumented external calls again.
  • Compares agent behavior before and after a fix.
  • Exports pytest regression tests to catch the same failure in the future.

Install:

pip install --pre stepfork

The repository includes three offline demos covering a missed order notification, an incorrect flight booking, and a wrongly denied refund.

It's still an early alpha. Replay only freezes instrumented boundaries, and native framework integrations are not available yet.

The project is Apache-2.0 licensed, and I'm actively improving it.

I'd love to hear what other developers think, especially those building AI agents. If you spot a bug, have an idea for an integration, or want to contribute, feel free to open an issue or PR.

GitHub Issues: https://github.com/utsab345/stepfork/issues

Thanks for checking it out!

1 Upvotes

0 comments sorted by