The phrase "run your own AI girlfriend" makes it sound like one program. In practice it is three pieces that talk to each other, and understanding which piece does what saves most of the frustration.
The three pieces

What talks to what in a typical local setup.
The front-end is what you see: chat windows, characters, avatars, settings. SillyTavern is the standard choice for character chat. It runs in your browser, stores characters and chats as files on your disk, and can connect to almost any back-end. It does not generate text itself.
The back-end loads the model and generates replies. KoboldCpp is a single program with no installation that is popular for roleplay because it exposes every generation setting. Ollama is the easiest way to get a model running in a few commands. LM Studio is a desktop app with a friendly model browser. Any of them serves the model on a local address that the front-end connects to.
The model is a large file, usually in the GGUF format, downloaded from Hugging Face. Its size in parameters - 8B, 12B, 24B - and its quantisation decide how good it is and what hardware it needs.
The hardware question
What decides whether a model runs well is graphics memory (VRAM). The model has to fit, plus room for the conversation itself. Quantisation shrinks models to fit: "Q4_K_M" is the common compromise between size and quality.
| Model size | Approx. memory at Q4 | Typical hardware | What to expect |
|---|---|---|---|
| 7-8B | 5-6 GB | 8 GB graphics card | Coherent, quick, forgets detail in long scenes |
| 12-14B | 9-10 GB | 12 GB card | Noticeably better characters and prose |
| 22-24B | 14-16 GB | 16-24 GB card | Close to good hosted apps for roleplay |
| 70B | 40 GB+ | Two cards or a high-memory Mac | Excellent, slow and expensive |
These figures are rough: longer context needs more memory on top. Apple Silicon Macs share memory between the processor and graphics, so a 32 GB MacBook can run models a 12 GB graphics card cannot. A model that does not fit spills onto the main processor and slows to a crawl - if replies arrive at a few words per second, drop to a smaller model before you try a harsher quantisation. A smaller model at good quality usually beats a bigger one squeezed too hard.
What you gain
- Privacy that does not depend on a policy. Nothing leaves the machine. No retention period, no training, no staff access, nothing to delete later.
- No filter beyond the model's own. You choose the model, including ones tuned specifically for roleplay. The legal floor on content still applies to you, but no operator is adding rules on top.
- No subscription and no credits. Generate as much as you like.
- Characters you own. Character cards are ordinary PNG images with the character sheet embedded in them. They move between front-ends and cannot be taken away by an update.
- Stability. A model that works today works the same way in five years. Nobody can swap it under you - the risk described in when your AI companion changes overnight.
What you give up
- Setup time. Expect an evening for the first working chat and weeks of tinkering if you enjoy it.
- Memory is your job. Hosted apps quietly summarise and store facts about you. Locally you manage this with SillyTavern's lorebooks, summaries and character notes - the same three mechanisms explained in how AI companion memory works, just with the controls exposed.
- Images and voice are separate projects. Picture generation needs its own model and usually more graphics memory; voice needs speech recognition and synthesis tools. All possible, none automatic.
- Mobile is awkward. You can open the front-end from your phone on home Wi-Fi. Using it away from home means exposing your computer to the internet, which needs care.
- Quality ceiling. The best hosted models are larger than anything most people can run at home, particularly for long-term consistency.
The middle route, and its catch
Many people run SillyTavern on their own machine but connect it to a paid cloud API instead of a local model. You get the front-end's control and character files, much larger models, and pay-per-message pricing that can be cheap for light use.
It is not private, though. Every message goes to the API provider, under its logging, retention and content rules - often stricter about roleplay content than dedicated companion apps. It is a different trade-off, not a local setup.
Safety basics
- Download software only from the official project pages on GitHub or the project's own site.
- Download models from well-known uploaders on Hugging Face, and prefer GGUF files, which contain model data rather than code.
- Read the model's licence. Some restrict commercial use; a few restrict certain content.
- Keep SillyTavern bound to your own machine or home network unless you have set up a password and understand what you are exposing.

SillyTavern is developed in the open on GitHub.
Who it is for
Run a companion locally if privacy is your main concern, you already own capable hardware, or you enjoy configuring things as much as using them. Stay with a hosted app if you want voice, pictures and long memory working out of the box, or mostly chat on your phone. Our guide to choosing a first app covers that route. Plenty of people end up with both: a hosted app for convenience, and a local setup for the conversations they would rather not store anywhere else.