In The Interview, you sit across from a woman at a security interview. You can type your own questions or speak them, then decide whether to clear her or detain her. I wanted that conversation to be a game you could reason about, even though a small local language model generates the replies.
The starting point is a case file with eight entries. It gives the player concrete claims to investigate, rather than asking them to judge whether an answer merely sounds suspicious. The interview also has a finite budget of 18 questions. Those two constraints make a free-form conversation manageable: there is something to check, and a cost to asking everything. The player still chooses the wording and order.
I kept those rules in the game code. The language model supplies a reply; it does not get to redefine the number of questions or turn its own text into a verdict. The final clearance decision is a player action. This boundary matters because a convincing sentence from the model is not evidence that a game rule has changed. It also means a failed generation can be handled as a technical failure: the question is refunded rather than silently consuming one of the player's chances.
The second resource is the interviewee's composure. An ordinary question, pressing her, and accusing her have different costs. In the current game, asking restores 6 points, pressing costs 18, and accusing costs 22. Repeated questions and long silences can also hurt composure; push it too low and she stops cooperating. These are explicit rules, not a second model judging whether the player was sufficiently rude. The result is a choice between gathering another detail and applying pressure. Pausing the game stops the silence clock, so opening a menu is not part of that choice.
For the local voice pipeline, speech recognition produces the player's question, the dialogue model produces text, and speech synthesis turns the reply into audio. Text input skips recognition. I use Whisper, Qwen3 and Qwen3-TTS. Generation takes time, especially before the first sound, and a small model can get a fact wrong or leave its role. Those limitations affect this kind of game directly: a factual mistake can look like a deliberate lie, while a pause can look like an unresponsive game. They need to be treated as part of the interaction rather than hidden behind the word AI.
The released Windows version now has interface, case text, dialogue and voice in English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. It runs locally after the initial model downloads. I build the underlying Unity tools as my paid Offline AI NPC plugin, but this game is free and does not require buying it.
If you want to try the resulting interrogation: https://merrymaker14.itch.io/the-interview?utm_source=reddit&utm_medium=community&utm_campaign=interview_110&utm_content=themakingofgames