Blog

How to Self-Host a Personal AI Assistant

August 10, 2026 · Privacy & self-hosting · 11 min

By , founder of Raegan

A small home server next to a laptop running a self-hosted personal AI assistant, suggesting infrastructure an owner controls.

To self-host a personal AI assistant you run an open model on hardware you control, put a chat interface in front of it, and connect your own tools. The path has six steps: decide between local hardware and a cloud server, pick a model and a runtime like Ollama or LM Studio, set up an interface such as Open WebUI, connect email and calendar through the Model Context Protocol, secure the setup, and keep it patched. None of it is exotic. All of it is ongoing.

TL;DR: A self-hosted AI assistant runs an open-weight model on your own hardware or rented server, so no business data enters a shared cloud pool. The model itself is now cheap to run: Stanford's 2025 AI Index found the cost to run a GPT-3.5-quality model fell from $20.00 to $0.07 per million tokens between late 2022 and late 2024, a more than 280-fold drop. The work that remains is operational: hardware, updates, security, and your time. This guide walks the six steps, then names the honest shortcut for owners who want the privacy without the second job.

This is a build guide for owners who want a private assistant and are willing to run it. It covers the real decisions in order, with current tools and honest effort estimates. If you would rather skip the maintenance entirely, the last section covers the managed-private route.

Step 1: Decide between local hardware and a cloud server

Start by choosing where the assistant lives, because that decides everything after it. The two options are a machine in your office or a rented cloud server. Local hardware gives you full physical control and no recurring compute bill, but you buy the machine and you fix it. A cloud server, or VPS with a GPU, removes the hardware purchase and runs anywhere, but it is a monthly cost and the data sits on someone else's metal.

The deciding factor is memory, not raw speed. Running a language model is mainly a memory problem: the model's weights have to be held in memory before it can answer anything, and larger models need more. As a rough, public rule of thumb circulated in setup guides, a small 7-to-8-billion-parameter model needs on the order of 8 GB of memory at common quantization, mid-size models more, and large models a great deal more. Treat those as general guidance and confirm the exact file size for any model you choose against its official listing, since quantization and context length both move the number.

If you do not want to buy a machine, rented GPUs are inexpensive by the hour. Public marketplace pricing in 2026 listed on-demand data-center GPUs starting around $1.39 per hour for an A100-class card on RunPod, billed by the second, with cheaper consumer cards available. For an assistant that runs continuously, though, hourly compute adds up to a real monthly figure, so price it as a recurring cost, not a one-off.

Step 2: Pick a model and a runtime

Choose an open model and the software that runs it. The model is the brain; the runtime is the engine that loads it and answers prompts. The good news for quality: the 2025 AI Index reported the gap between the best open-weight and closed models narrowed from 8.04% to 1.70% on the Chatbot Arena leaderboard in a single year, so open models now handle everyday assistant work without an obvious penalty.

Two runtimes cover almost everyone:

Pick by taste: Ollama if you are comfortable in a terminal, LM Studio if you want a graphical app and built-in model search. Both pull from the same open-model ecosystem, so you are not locked into either.

The quality gap between open and closed models nearly closed in a year Column chart showing the performance gap between the best open-weight and closed AI models falling from 8.04 percent to 1.70 percent over one year, per Stanford HAI 2025 AI Index. Open vs closed model quality gap Percentage-point gap on the Chatbot Arena leaderboard, year over year 8% 4% 0% 8.04% Earlier year 1.70% One year later Source: Stanford HAI, 2025 AI Index Report.
Open-weight models closed most of the quality gap in twelve months, so a self-hosted assistant no longer means a worse assistant. Source: Stanford HAI, "2025 AI Index Report."

Step 3: Set up the interface

Put a proper chat interface in front of the model so it works like a real assistant, not a terminal. A runtime can answer prompts, but you want conversation history, document upload, and a usable screen on your phone. Open WebUI is the common choice here. It describes itself as an "extensible, feature-rich, and user-friendly self-hosted AI platform designed to operate entirely offline."

Open WebUI installs by Docker, pip, uv, or a desktop app, and its docs recommend Docker or Python for production deployments. It connects to Ollama and to any OpenAI-compatible API, which is exactly what LM Studio exposes, so it sits cleanly on top of either runtime you chose in Step 2. On first launch it asks you to create an admin account, and it supports multiple users with role-based access if you later add a teammate.

Two setup notes from the project's own guidance matter. First, persist your data: when running in Docker, mount the data volume or you can lose your database and chat history. Second, do not rush it onto the open internet. The next step is why.

Step 4: Connect email and calendar

Give the assistant access to your tools through the Model Context Protocol, the open standard for connecting AI applications to data and actions. MCP, introduced by Anthropic in late 2024, is described in its own documentation as "an open-source standard for connecting AI applications to external systems," likened to "a USB-C port for AI applications." An MCP server exposes a set of tools the assistant can call, like searching email or creating a calendar event.

For email and calendar specifically, Google now offers first-party MCP servers for Google Workspace, currently through its Developer Preview Program. Per Google's developer documentation, each Workspace product has its own server, and they let an AI agent "read data" such as "Search emails, retrieve files, and list calendar events" and "take action" such as "Create draft emails, upload files, and schedule meetings." Critically, the servers "inherit the same permissions and data governance controls as the user," so the assistant can only do what your account can.

Connect the minimum first. Read-only access to your calendar and inbox is enough to get a daily summary and draft replies, and it limits the damage if anything goes wrong. Which brings up Google's own warning, in that same documentation, that connected tools "can read, modify, and delete data in your Google Account," and that you should "only use trusted tools" and "review all actions." Treat write access, especially sending email, as a deliberate later step with a human approving each action.

Step 5: Secure it

Lock the assistant down before it touches the network, because a private model on an exposed server is no longer private. Open WebUI's documentation is explicit that you should not expose an instance publicly until accounts, HTTPS, backups, and API access are properly configured. The privacy you self-hosted for lives or dies on this step.

The instinct to keep data close is sound, but location alone is not security. In Cisco's 2025 Data Privacy Benchmark Study, 90% of organizations said local storage is inherently safer, yet 64% worried about exposing sensitive information through AI. Self-hosting removes the shared-pool risk; it hands you the patching, access control, and monitoring in exchange. The stakes are concrete: IBM's Cost of a Data Breach Report 2025 put the global average breach at $4.44 million and linked unsanctioned "shadow AI" to roughly one in five breaches.

A practical baseline for a single-owner setup:

Step 6: Maintain it

Plan for upkeep, because a self-hosted assistant is a small service you now operate. Models get new versions, the runtime and interface ship security updates, GPU drivers drift, and disks fill. Each model upgrade is a test-and-redeploy cycle, and someone has to notice when the box stops answering. For a solo owner, this is the real cost of self-hosting, and it is measured in attention rather than dollars.

The compute is genuinely cheap now. The expense moved to operations: keeping the stack patched, watching for failures, restoring from backup when something breaks, and re-validating integrations after each update. None of it is hard in isolation. It simply never ends, and it competes with running your business. That trade-off is the whole reason the next section exists.

Where the effort sits when you self-host an AI assistant Horizontal bars showing relative ongoing effort across six self-hosting steps, with security and maintenance carrying the most ongoing effort and choosing a model the least. Illustrative, based on the workflow described in this guide. Ongoing effort by step The setup is one-time; security and maintenance recur forever. Hardware vs VPS Model + runtime Interface Connect email/calendar Security Maintenance Illustrative, based on the six-step workflow in this guide. Darker bars recur indefinitely.
Setup is finite. Security and maintenance are the steps that never close, which is the honest cost of running it yourself.

Or skip the maintenance: managed-private

If you want the privacy of self-hosting without becoming a part-time systems administrator, a managed-private assistant runs the same kind of private setup for you. The privacy benefit owners actually want does not come from doing the work yourself. It comes from where the data sits and who it is pooled with. A managed-private assistant keeps your data on a server dedicated to you, with no shared training pool, while a provider handles the hardware, updates, security, and the integrations you would otherwise wire up by hand.

This is the lane Raegan sits in. It is built on Hermes, an open-source agent from Nous Research, runs on the customer's own server with the upkeep handled for them, and gates outbound email behind owner approval, so business data is never sold, shared, or dropped into a common cloud pool. It is one option among several. For owners who reach Step 6 and realize maintenance is the part they do not want, it removes the second job while keeping the data private.

To weigh the two paths directly, see self-hosted vs managed-private AI assistant. For the privacy fundamentals, start with what is a private AI assistant, and for help choosing a model, open source AI assistants explained covers the trade-offs. The self-hosting question also sits inside the broader AI personal assistant category, if you are still deciding what kind of assistant you want at all.

Frequently asked questions

What do I need to self-host a personal AI assistant?

You need somewhere to run it, an open model, a runtime, and an interface. That means either a machine with enough memory or a rented GPU server, a model from a library like Ollama's, a runtime such as Ollama or LM Studio, and an interface like Open WebUI. Then you connect your tools and secure it.

How much memory do I need to run a local AI model?

Enough to hold the model in memory, which scales with model size. A small 7-to-8-billion-parameter model needs roughly 8 GB at common quantization per public setup guides, with larger models needing more. These are general estimates; quantization and context length change the figure, so confirm the exact file size against the model's official listing before buying hardware.

Can a self-hosted assistant access my email and calendar?

Yes, through the Model Context Protocol. Google offers first-party MCP servers for Google Workspace, currently in developer preview, that let an assistant search email, list calendar events, draft replies, and schedule meetings. They inherit your account's permissions. Start with read-only access and require approval before the assistant sends anything.

Is self-hosting an AI assistant secure?

It can be the strongest option, but only if you do the security work. Self-hosting removes the shared-cloud-pool risk that worries most owners, which is why 90% of organizations told Cisco in 2025 that local storage feels safer. It also hands you patching, access control, and backups. A neglected self-hosted server can be less safe than a well-run managed one.

Is it cheaper to self-host than to use a cloud AI tool?

The compute is cheap; the operations are not. Stanford's 2025 AI Index found inference for a GPT-3.5-quality model fell over 280-fold to about $0.07 per million tokens. For a single owner, the deciding cost is rarely the model bill. It is the hardware, security work, and hours spent keeping the assistant running and updated.

Do I need a powerful GPU, or can I rent one?

Either works. A capable GPU or a Mac with generous unified memory can run mid-size models locally, but you can also rent a data-center GPU by the hour. Public marketplace pricing in 2026 listed A100-class cards from around $1.39 per hour on RunPod. For an always-on assistant, price rented compute as a recurring monthly cost.

Sources

  1. Stanford Institute for Human-Centered AI. "2025 AI Index Report," 2025. https://hai.stanford.edu/ai-index/2025-ai-index-report
  2. Stanford Institute for Human-Centered AI. "2025 AI Index Report, Research and Development," 2025. https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
  3. Ollama. "Ollama (GitHub repository and model library)," 2026. https://github.com/ollama/ollama
  4. LM Studio. "LM Studio Documentation," 2026. https://lmstudio.ai/docs
  5. Open WebUI. "Open WebUI Documentation," 2026. https://docs.openwebui.com/
  6. Model Context Protocol. "What is the Model Context Protocol (MCP)?," 2025. https://modelcontextprotocol.io/introduction
  7. Google for Developers. "Configure the Google Workspace MCP servers," 2025. https://developers.google.com/workspace/guides/configure-mcp-servers
  8. Cisco. "2025 Data Privacy Benchmark Study," 2025. https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2025/m04/cisco-2025-data-privacy-benchmark-study-privacy-landscape-grows-increasingly-complex-in-the-age-of-ai.html
  9. IBM. "Cost of a Data Breach Report 2025," 2025. https://www.ibm.com/reports/data-breach
  10. RunPod. "Cloud GPUs," 2026. https://www.runpod.io/product/cloud-gpus

Raegan is a private, self-hosted AI assistant that runs on your own server, with the upkeep handled for you. Get early access.

Meet Raegan, your private AI chief of staff

She handles email, notes, research, calendar, and a daily briefing - private and self-hosted. Join the early-access waitlist.

Get early access