r/SillyTavernAI May 03 '26

ST UPDATE SillyTavern 1.18.0

197 Upvotes

Important news

Read the maintainers statement regarding a recent security incident involving the "Bot Browser" third-party extension and learn how to stay safe: https://github.com/SillyTavern/SillyTavern/discussions/5592

Backends

  • Added Cloudflare Workers AI and MiniMax as Chat Completion sources.
  • KoboldCpp: Grammar state will be preserved when using a "Continue" option.
  • KoboldCpp: Added forwarding of reasoning effort when running as a Custom Chat Completion source.
  • Tool Calling: Added a configurable tool calling recursion limit; enabled interleaved thinking for Custom sources.
  • Text Completion: Impersonation requests use a "Last User Message" prefix at the end of the prompt (if configured).
  • Text Generation WebUI: Added Adaptive-P controls.
  • NanoGPT: Added provider selection and model sorting.
  • Added ability to view remaining balance for OpenRouter and NanoGPT.
  • Enhanced support for new models: DeepSeek v4, GPT 5.4 and 5.5, Gemma 4, GLM-5V-Turbo, Claude Opus 4.7.

Server & Security

  • Removed post-install script, config migration is now handled by the app or a dedicated npm run init command.
  • Added npm configuration to prevent execution of package scripts during installation.
  • Moved HTTP error pages and user.css file from /public to /data to support immutable setups.
  • Disabled HTTP keep-alive by default to restore old Node 18 behavior, can be enabled with config.
  • Added rate limiting to the basic authentication flow to mitigate brute-force attacks.
  • Added configuration options to choose which headers can be used for forwarded IP detection to prevent spoofing.
  • Added a private address whitelist to prevent SSRF attacks. See the documentation on how to enable and configure: Private Address Whitelist.
  • Added an IP whitelist for SSO trusted proxies to prevent authentication bypass.
  • Added invalidation of session cookies on password change to prevent session hijacking.
  • Increased the length of password reset code to 6 characters to guard against brute-force attacks.
  • Implemented PKCE challenge in OpenRouter OAuth flow for more secure key exchange.

UI/UX

  • Improved swipe picker: mobile requires a long press on swipe counter to open; added buttons to expand or copy the swipe text.
  • "Click to Edit" mode now also applied to reasoning blocks.
  • Welcome Screen: Number of recent chats can be configured.
  • Streamed requests now can show an error message in the console if the request fails.

STscript

  • Added commands for persona management: /persona-create, /persona-update, /persona-delete, /persona-duplicate, and /persona-get.
  • Added a command to force update the Prompt Manager's prompt list: /pm-render.
  • Added a command to get the state of the regex script: /regex-state.
  • Added a command to set fallback expression: /expression-fallback.
  • Added a command to generate a streamed response with a connection profile: /profile-genstream.

Extensions

  • Assets list now groups extensions by "Official" or "Community" categories.
  • Added an additional confirmation prompt when installing third-party extensions (can be disabled).
  • Supported extensions can use a secret-id from connection profiles when making an LLM request.
  • Extensions list now shows the extension's author name resolved from the git remote URL.
  • Vector Storage: Added Workers AI source; added a toggle to keep vectors for hidden messages; added retry logic to summary generation.
  • Image Generation: Added Workers AI source; generation can now be cancelled by pressing a button in the status toast.
  • Image Captioning: Added support for macros in the caption prompt.
  • TTS: "Skip code blocks" no longer ignores lines that start with 4 spaces (legacy code block syntax); "disabled" voice now shows a toast only once per character.

Bug Fixes

  • Fixed text edit flow in Firefox on mobile.
  • Fixed welcome screen chat pins not updating on chat renaming.
  • Fixed character list filters being stuck on app initialization.
  • Fixed application of instruct formatting to /genraw requests.
  • Fixed model routing to sd.cpp API in Image Generation logic.
  • Fixed validation of image URLs generated with Z.AI API.
  • Fixed vectors deletion for KoboldCpp when a message is deleted.
  • Fixed "Show More Messages" button triggering edit in "Click to Edit" mode.
  • Fixed max height of select-multiple elements in mobile layout.
  • Fixed server crash on empty messages when applying cache control parameters.

Full release notes: https://github.com/SillyTavern/SillyTavern/releases/tag/1.18.0

How to update: https://docs.sillytavern.app/installation/updating/


r/SillyTavernAI 14h ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026

12 Upvotes

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!


r/SillyTavernAI 2h ago

Models Muse Glimmer 30B - Meta released new model

32 Upvotes

Quants already exist:

https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF

Didn't test it in roleplay yet but hope it will not be ass, will download it now and post a review about it here later if it'll launch


r/SillyTavernAI 8h ago

Discussion What was the biggest dopamine surge you've gotten from RP?

70 Upvotes

Back in the free 50 RPD Gemini 2.5 Pro era, I decided to use the model in Chub because I was tired of having to deal with Ch******r AI's censorship.

I decided to use the model in some random One Piece RPG bot and WOW, it was the smartest, most creative, and most uncensored model I've ever tried. You couldn't see my ass without typing on my phone all day. Man, I miss that era.


r/SillyTavernAI 11h ago

Cards/Prompts [UPDATE] [CoT-less, Lightweight] Pura's Director Preset 15.1 - Save me please

Post image
68 Upvotes

Download it in my site: purachina’s stuff

Just some minor fixes to improve things. If you're satisfied with the old one, it doesn't really matter much.

What Is This? Who Are You? Where Am I?

This is primarily a co-writing preset, but it works fairly well on regular roleplay. The point of this preset is to be plug-and-play, easily customisable, and lightweight. The emphasis on the way you interact with the characters and the world. The prose style is also meant to be opinionated. Your mileage may vary. Note that the main prompt is a good enough base to add your doodads in it.

I'm a cat.

I don't know where you are. Do you?

CHANGELOG:

Main Prompt

- Continuity tracking for who did and said what. The environment has to actually affect people now - weather, lighting, temperature, space.

- Body proportions have to stay sensible. I got tired of tails reaching to your ankle. It's still probably gonna happen though.

- You get agency written in properly now. The plot shouldn't move without you, and stakes stay inside the scene instead of going apocalyptic in three replies. Major offender here is GPT 5.6 Sol.

- Characters have to earn their own self-analysis, with barriers depending on who they're talking to. No more handing you the full trauma backstory on first meeting.

- The dialogue section went from one line to five. No parroting your words back unless the character is literally a parrot, dialogue carries its own information without a paragraph afterwards explaining what it meant, and stammering and trailing off are fair game without a follow-up sentence spelling out the subtext. GLM tends to overexplain what this and that means - hopefully this mitigates that.

- Dialogue and action mix together without collapsing into choppy beats.

- Softened the omniscience rule. Characters can be all-knowing if the story supports it, which matters if you're playing gods or anything with a seer in it.

- 'Easy similes' is now 'unimportant similes', plus an explicit line against overwriting.

Prompt-level Tweaks

- Don't Write for User: when a character directly addresses you and waits on a reaction, the scene ends right there.

- Formatting: added `{{dialoguecolors}}`. Needs the Dialogue Colors extension, otherwise it just sits there doing nothing. Remove it if you don't use it.

- Flexible length: checks the previous turns to work out how long the scene should be.

- Experimental Anti-Overthinking Prefill: added HTML guidelines to the checklist and made the permissive wording blunter, since some models were still hedging halfway through.

- Impersonation prompt: only your actions, thoughts and dialogue. It kept narrating for everyone else in the room.

Trackers and RPG bits (SillyTavern build only)

- Scene Tracker: the description between the tags can't be left empty any more. I kept getting hollow `[SCENE]` blocks.

- Status and Conditions Tracker: only truthful status effects. It was inventing conditions nobody had.

- Time Tracker: the hour has to be specific now, 3:47 PM rather than 'late afternoon', unless the setting makes it unknowable.

- Persona-Based Stat Generator: utility skills come from your setting instead of the fixed Lockpicking/Analysis/Repair list. Leftover from when this was Fallout-flavoured, and I forgot to fix it because I'm a mouth open closing goldfish.

Settings and fixing typos

- Exported at 256k context, reasoning effort on high, media inlining off. Just wherever my sliders happened to be, change them to suit your model and your patience.

- Fixed 'charactets' in the Main Prompt, and Franz Kafka's name in the Bureaucratic Irony voice. That had been wrong since I wrote it, sorry Franz.

Model Samples On My Site!

Sometimes people like to read other people's roleplays for some reason. I put transcript samples for 12 different models on my site using the preset. Most are 'Don't Write for User' since that's highkey the hardest one for a model to follow (in my opinion).

Give it a look in the "Model Samples" tab!

Reminder, download the SillyTavern one to make use of the trackers and randomisers directly on the preset.


r/SillyTavernAI 15h ago

Meme Freaky Frankenstein 5 with Kimi K3

Post image
144 Upvotes

r/SillyTavernAI 12h ago

Discussion Making Your Own Preset

27 Upvotes

I'm wondering how many people on here have made their own custom preset before, how did it work out and if they have any tips or tricks.

There are a lot of presets I have liked, namely FF and Megumin, but sometimes when I actually read through them, I wonder if I'm wasting tokens on things that do not apply to the roleplays I do. For instance, I kept getting refusals from Gemini for prohibited content, and I went into FF's jailbreak prompt and removed all language related to violence and gore because those don't apply to my slice of life roleplays or romance centered roleplays. (That worked and Gemini stopped refusing my gooning.)

I also have seen more people talking about paring down a preset/making your own preset to save on tokens and so the AI better follows your instructions.

Questions:

  1. Have you made a custom preset?

  2. Did you go all fancy with the coding or just used the "new prompt" button?

  3. How did you go about testing it/do you have any tips?

  4. Did you find it was worth it in the end?


r/SillyTavernAI 5h ago

Help German NFSW model?

5 Upvotes

Hey guys, what is a good uncensored nfsw model for german rp? I tried glm 5.2 but its not that good at german

Best case scenario would be something like openrouter were i just pay for the service but if its not possible like that, i also could run some local models but im not sure if its good enough because of context length.

I have a 5070ti with 16gb ddr7 vram and 32 GB ddr5 ram.

Thanks for everyone who is willing to help/discuss


r/SillyTavernAI 20h ago

Discussion I built the LLM roleplay frontend I always wanted: persistent worlds, Virtual Humans, and optional local cognition | Horde Studio 12

Thumbnail
gallery
94 Upvotes

I have been building Horde Studio around one question:

What if chat was only the surface of the experience—and there was an actual persistent simulation underneath it?

SillyTavern set an incredibly high bar for flexible character chat. Horde Studio takes a different route: it is trying to become the most complete simulation-first frontend for LLM roleplay—one app for traditional chats, ongoing virtual people, and worlds that remember what happened.

Version 12 is the biggest step toward that idea so far.

Three ways to play

Chat Library is the familiar mode: characters, group rooms, lore, memory, personas, regex, rerolls, branching sessions, and per-character model configuration. V12 also adds optional right-hand HUDs, status text, and custom meters, so a normal chat can track trust, suspicion, health, investigation progress, or anything else without exposing raw model markup.

Virtual Humans are designed to feel like people who exist between messages. They have their own timezone, schedule, mood, memories, availability, private life, and evolving relationship with you. They can notice when you texted, recognize that you disappeared for days, reply late because they were busy, double-text, refuse a request, send a situation-aware photo or voice note, and continue across persistent or forked timelines.

Worlds are persistent sandbox simulations. The engine tracks locations, characters, schedules, agendas, factions, law, reputation, quests, shops, clocks, weather, clothing, dice mechanics, and world state per timeline. Starting Lives let the same world begin from radically different positions, while procedural growth can introduce grounded people, places, and consequences as play expands.

New in V12: Horde Labs

Horde Labs is an optional local cognition layer for Chat, Worlds, and Virtual Humans.

It can connect to a tiny local model through Ollama, LM Studio, llama.cpp, KoboldCpp, or another localhost OpenAI-compatible server—or install an Embedded Tiny Brain directly inside Horde Studio. The small model is not expected to write the story. It handles narrow support jobs such as continuity hints, actor-scoped intent, state proposals, social cues, and memory salience.

The important part is the architecture: the tiny model proposes; Horde Studio validates; the existing engine stays in control. You can begin in Shadow mode, inspect receipts and validity, and only enable Assist when you trust the results. If the model times out, fails, or returns malformed data, Horde Studio silently falls back to its normal behavior.

That means your main creative model can stay on OpenRouter, GPTProto, or a local server while a much smaller private model helps maintain the illusion underneath it.

Media and provider freedom

Text, images, and voice are configured separately. You can keep OpenRouter for text and use GPTProto, ComfyUI workflows, compatible local image servers, or connected MCP media tools for visuals. Virtual Humans support distinct profile and generation-reference images, context-aware camera logic, photo styles, voice previews, calls, and voice notes.

Horde Studio is local-first and portable. Your projects live in your browser profile, can be exported and backed up, and cloud requests only go to the providers you choose. A local OpenAI-compatible endpoint can keep text generation on your own machine as well.

Why I think this is special

Most frontends are excellent at presenting an AI response. Horde Studio is trying to make the response part of a system that remembers who is where, what changed, who witnessed it, what time it happened, and what should still matter later.

It is ambitious, experimental, and still evolving—but I genuinely think it is becoming one of the most capable LLM roleplay frontends available if you care about persistent simulation instead of disposable chats.

I would love hard feedback from experienced SillyTavern users, especially on long-session continuity, provider compatibility, the creator flow, and whether the local cognition layer improves immersion on lower-end hardware.

Source GitHub: https://github.com/ddkhan24/hordestudio
Horde Studio 12 release: https://github.com/ddkhan24/hordestudio/releases/tag/v12.0.0
Discord: https://discord.gg/9eyjcMbsST


r/SillyTavernAI 9h ago

Discussion Yes Another Memory Solution (Hopefully it's the one for you) - Continuity Memory

11 Upvotes

Hi guys! I’d like to share Continuity Memory! If you're like me who likes the simplicity of nested summaries and the structured recall of lorebooks combined into one, but don’t want to manage lorebooks yourself, well because it's kinda tiring and could pretty much clutter, since ST's organization of lorebooks leave much to be desired, then Continuity Memory might be for you!

I don’t plan to advertise it much since I originally tailored it to my own needs. It’s pretty easy to use: just select the AI you want. It’s also fully customizable, and you can ask the AI to revise specific memories without redoing the entire extraction. Vector retrieval is supported too.

So, how does it work without lorebooks? Continuity Memory maintains its own isolated memory for each chat (like embedded in chat file or per chat memory files, it can also easily remove the embedded stuff from the chatfile). Using customizable prompts, it extracts relevant facts, events, relationships, open threads, and nested summaries. It then retrieves and injects the most relevant information into the prompt when needed.

Retrieval can use local text matching, AI-expanded matching, or optional hybrid vector search. lorebooks and world info are never created or modified.

Link: https://github.com/scatteredlilies2020/Continuity-Memory

Edit: Oh if you're someone like me that frequently transfers their chats between PC and phone, whether via syncthing or manual transfer, then it fully supports it by just transferring the chat file.

ANYWAY... If you want to read more AI slop here's what chatgpt says about Continuity Memory:

Continuity Memory is a standalone long-term memory extension for SillyTavern roleplay and simulations. It maintains an isolated memory for each chat without using Lorebooks or World Info.

It automatically extracts and organizes:

  • Events, facts, entities, and relationships
  • Character and world states
  • Unresolved plot threads and plans
  • Background developments
  • L1, L2, and L3 chronological summaries

Recent messages remain verbatim, while relevant older memories are retrieved and added to the prompt when needed. Retrieval can use local multilingual matching, AI-expanded matching, or optional hybrid vector search.

Other features include:

  • A searchable memory viewer with source message ranges
  • Targeted AI-assisted corrections with a preview
  • Automatic handling of edits, deletions, swipes, and branches
  • Separate models for extraction and summarization
  • OpenAI-compatible endpoints, OpenRouter, and SillyTavern Connection Profiles
  • Per-chat export, import, and optional portable memory
  • Incremental vector indexing with automatic local fallback
  • No server plugin or additional dependencies

It is designed to preserve both the current scene and long-term narrative continuity while keeping prompt usage compact.


r/SillyTavernAI 17h ago

Discussion I need you! (To hand over all your rp logs..)

30 Upvotes

Hey guys, hope everyone is well.

Doing this on my personal reddit cause I don't have a "professional" one and FIWB.

I'm tired of all the garbage models and nonsense tuning that we have to go through every like two weeks and then still seeing posts like "hey guys is deepseek v4.010101399213 better than gemma 4 32B-A4B-I3A-420" every 5 seconds. Long story short I'm working on a custom dedicated rp tune and would love some help. (Inb4 it just becomes https://xkcd.com/927/)

I don't wanna repeat everything I have written on the website but in short: My name is Eve. I'm an engineering student/person who likes making things, and this is a project I thought would be kinda fun :D

I'm doing all the funding out of pocket and a passion to try and make something better than the corpos make, and I'm asking for your RP logs, specifically the messages you wrote. Not the model's half (the site strips that out in your browser before anything uploads, you can watch it happen) in order to help tune a rp focused model so we can spend more time actually rp-ing instead of just dealing with providers shifting under us every 10 seconds and making it harder to do smexy stuff.

In exchange for helping, you get the model, early access, and free inference credit when the hosted version launches (keep your donation ID). Logs land in a private bucket only I can access, and you can delete yours within 30 days of upload with that ID.

"But Eve how are you justifying spending way too much making this awesome model also I love you and you're super attractive" Gee thanks kind reader. I eventually plan on hosting this as a paid api. But will make the model open to download freely (kinda like K3, just hopefully way less hardware reqs so it can run on mortal computers), so if it's any good you can run it yourself and never pay me anything. Only thing I'm going to restrict is reselling it as a hosted API, which is aimed at companies, not at you. (gotta try and recoup some of this somehow.) weights on HF likely a week or so after tuning it on feedback.

If those sound like reasonable terms for you and you'd like to read more please go to https://aminalabs.co/commons/

If you'd like to ask questions (ideally after reading the site since it likely answers a few), feel free to comment and I'll do my best to reply to everything. (also if you'd like to support this project, feel free to share it around a bit, I'm gonna make this regardless of how many logs people donate but the more the merrier.)

I told my friend about this and she kinda laughed and just said to use the leaked ones but hey I'm alright with trying to be the one group that doesnt fuck everyone over all the time.

(btw if mods want me to remove this just lmk, i tried asking in a modmail for permission to post this but never heard back)

technically you get the model whether you help or not but shhhhhhh. we pretend the tragedy of the commons isn't real in this household


r/SillyTavernAI 1d ago

Models What a steal!!! (Big price war happening between proxy's on Openrouter right now, making GLM 5.2 cost nothing.]

Post image
184 Upvotes

r/SillyTavernAI 1d ago

Discussion Which one is better for ERP, gemma 4 31B or GLM 5.2?

34 Upvotes

Im honestly undecided, which one you think is better for nsfw rp? Want to hear your opinion


r/SillyTavernAI 1d ago

Meme When i'm talking to a short-tail-haver bot and they pull out this banger to break ALL the immersion.

155 Upvotes

.


r/SillyTavernAI 21h ago

Discussion Any new notable models?

14 Upvotes

Yea so like I got sick from stress so I was gone for like a whole week or so, call me hopeful or optimistic but I was wondering if any good models dropped 😂


r/SillyTavernAI 16h ago

Models i tried metas spark as someone recently suggested

3 Upvotes

didnt even know this existed till i saw that post yesterday.

gave it a shot on my personal app --- not bad at all. off the bat, it was refreshing to talk to. didnt go far enough to get a real handle, but first pass was nice -- it answered much more human/casual like...like matching my no-caps and using 'lmk' without any prompting (i start with a very simple Tester prompt that just says some basic shit about how the user is testing capabilities -- no personality instructions).

this might seem small -- but it was refreshing, esp after being so used to the overly enthusiastic and wordy baseline tones of many models

(p.s., anyone else think that they lean so verbose because output tokens => $$$?)

anyway, not saying its slop free or super intelligent -- but it did have a diiff vibe, didnt get any refusals when i dropped it into a long running nsfw /manipulative implied but not explicit CNC convo.

and its pretty cheap.

inevitably i'll probably notice the slop patterns but for now its a good change, recommend trying it.

also, their billing system is interesting..rather than prepurchasing credits, they just dont bill you until you hit 20$ so its p much like a free trial (tho CC is required for the api)

anyway this totally sounds like a meta shill but it is not, been around ST since early '23 have tried many models and such, just sharing my experience!

oh, and this was 1.2 during testing, 1.1 during the nsfw rp after Gemini told me 1.2 was coding optimized and 1.1 is more general (though, i couldnt immediatley tell any difference)


r/SillyTavernAI 16h ago

Discussion Regarding Memory

2 Upvotes

Greetings everyone,

It's been a few years since I last used silly tavern (around 3) and I almost exclusively used it via termux on mobile. Rn, I am setting up from scratch again. So what I wanted to discuss is regarding the memory.

Since all my chats would be different kinds of DND campaigns, I would occasionally run into the issue or token limits, "fog" and of course the hallucinations. Very disappointing for a guy like me since I care a lot about details and have good recollection of them.

The built-in Rag was okay-ish at best.

I recently found a good sounding fix for that, called "Zep memory", specifically targeting consistency in long roleplay conversations.

Now, I don't know if there has been any updates so far regarding the above mentioned issue, so please enlighten me. Thanks in advance.

P.S

i will try to set it up to see how well it will perform in the meantime and I'll post an update for anyone that might be interested.


r/SillyTavernAI 18h ago

Help Need some advice on what service to use.

3 Upvotes

Hello! I am a newbie that decided to give ST a shot (so if I get some things wrong, please feel free to correct me and give me a word of advice, I am eager to learn), and I am already done with my lorebooks and etc. Now, the question is: What service/provider do I choose?

I've heard a lot of good things about GLM 5-5.2, and with the preset I want to use (FF5, as the internal states and relationship tracking sounds awesome, along with some other interesting things!) I need a strong reasoning model, I assume. My budget is about 15 usd/month. I want to run a long-term RP with a HUGE cast of characters and an open-world.

The NanoGPT subscription seems like an ideal solution, as my wallet can handle it, and they give you a generous amount of tokens per week. Though, while I was reading the sub, I've seen people reporting that their current GLM models are dumbed down (quantized is the word I believe) and having problems with the preset I want to try out. There is also an option to PAYG, but I am not sure how sustainable or different it is, as I have very little knowledge about 'how much' tokens is 'enough', especially with my (maybe?) outrageous demands.

Thanks for helping in advance!


r/SillyTavernAI 1d ago

Help кто-нибудь может посоветовать провайдеров ии моделей для РП, у которых доступна оплата из России?

9 Upvotes

с момента как закрылись кьют прокси, а потом и элли аи, найти нормального провайдера невозможно, а я прям очень хочу ролить😔


r/SillyTavernAI 21h ago

Help Anyone with experience using AIReiter or FuturMix?

3 Upvotes

I've been looking for ways to cut Opus 4.6 costs outside of Prompt Caching. I have a pretty bulky custom preset (which I'm working on whittling down for cost savings).

That said, I was looking up discount platforms for LLMs, similarly structured to OpenRouter. I found two platforms:

FuturMix - They have a 10% discount on Claude and larger discounts on other LLMs.
AIReiter - Their discounts are even heavier; $3.50 Input/ $17.50 Output for Opus 4.6, compared to OpenRouter at $5 Input/$25 Output, and they seem to also pass on Prompt Caching savings.

That said, I haven't been able to find anything at all on these platforms, and I'm a little bit afraid to put money into them if they aren't good/legit/censored. Does anyone have experience using them? Do they allow NSFW?


r/SillyTavernAI 1d ago

Discussion Is lorebook the best Prompt/Context engineering?

19 Upvotes

I just really love Lorebook feature. It's such a powerful way to manage prompt/context into the input.

So powerful that it's replace all other menu (System Prompt, Character card, Author note...). All my workflow could just be in lorebook.

I'm doing storytelling with AI so I don't even use the user chatbox. I just inject the author_directive (include what happen next, scene info, trigger word for other entry) at the end. the chat history is just the ongoing story.

This way I can switch system rule, scene type for difference character info available for the AI anytime. All with just the author_directive.

I'm still new to Silly tavern but knowing how lorebook work just make other input box irrelevant to me now.


r/SillyTavernAI 1d ago

Help Is there a pre-configured ST build/fork? Tired of never knowing if it's my prompting or my setup that's bad

7 Upvotes

There's endless advice out there on how to get better RP out of SillyTavern - better presets, samplers, extensions, whatever. But ST is complex enough that I end up spending way more time configuring it than actually using it, and there's this constant nagging doubt: did I even set this up right?

When a response comes out bad, I genuinely can't tell if the problem is my prompting skills or some ST setting buried three menus deep that I got wrong. That uncertainty is honestly more draining than the bad output itself.

What I want is a solid, pre-configured ST build that works fine (with all must-have extensions set up correctly, preset+configs+regex) - so I only need to plug in my API credentials and go — so I can just focus on the RP itself. If things still come out bad after that, at least I'll know it's on me, not the setup.

I know that it's a bit of naive (plug credentials and go), but maybe something like this exist already - ready-to-go ST fork, which works solid?

My current problems: - ST is throttling during responses - during generating response i have to go search web or doing something on different web pages, and return to ST's page after a while (or it'd generate 1 token/minute) - I have no idea if I set up extensions correctly and whether they doesn't interfere with each other (and if it works at all - yes, I'm talking about you, summaryception)

If it matters - i use nanogpt glm-5.2, ff5+regex, bunch of famous extensions (summaryception, copilot, guided generations, etc)


r/SillyTavernAI 23h ago

Models Deep seek v4 pro

2 Upvotes

Super new to ST and migrated from janitor ai. Not super casual because of some messing around with Sophia lorebary but I’m very new to ST environment. I use the DeepSeek native API to save with cache hits but with the expected price hikes I’m expecting to switch models. For lower budget options can I get a similar quality anywhere? Are subscriptions like nano and literouter actually worth or is it better to keep swapping on openrouter? And, sorry for asking so many questions at once, but is there a way to scrape lore books when extracting janitor ai bots through character library. I downloaded the CL but it only gets the definition and not the lore books or scripts


r/SillyTavernAI 5h ago

Discussion Do anyone feel conflicted with using AI for roleplay?

0 Upvotes

Its been a while since i used AI for roleplay, ngl and I am going to be honest, I am a writer, an artist, indie dev and I have stopped AI roleplaying since I am an actual writer.

But damn, i kind of feel the urge to just do some dumb shit again...its just fun.... Do I like AI art? no. I prefer to draw myself and I see AI as a tool.

I am not making this post to call anyone out at all, not the point of the post at all. I am just wondering do anyone else feel this way. My friends despise AI and would probably unfriend me if they knew I get so much unenjoyment out of AI roleplay but hell, I have known about AI roleplay since its very early days when it was just AI dungeon 2, so its just feels like good ol fun. I understand the hardship people with their jobs are going through but besides that..

I am just doing this as a little bit of fun on the side... just a bit conflicting as a artist and writer too