Acutis Gate logo Acutis Gate identity for AI at work

A local LLM for a small town or school district

Short answer: one small server or mini-PC can run a capable open model for a town hall or district office: 16 GB of RAM runs a 7-billion-parameter model on CPU at a few words per second, enough for drafting, summaries and questions about your own documents. Put Open WebUI in front of it for staff, and connect it to your systems through a gateway that runs every action as the person who asked, so residents' and employees' data never leaves the building.

Sizing

  • 8 GB RAM: a 3B model (for example qwen2.5:3b). Fine for short drafting and simple questions.
  • 16 GB RAM: a 7B model (qwen2.5:7b). The sweet spot on CPU; expect roughly 3 to 8 words per second.
  • 32 GB or more: a 14B model (qwen2.5:14b). Better answers, slower on CPU.
  • A GPU (or an NVIDIA DGX Spark class box) makes it fast enough for a whole department at once.

Pick a model that supports tool calling if you want it to work with your systems, not only chat.

A working setup

  1. On Windows or a Mac, install the Ollama app from ollama.com/download. It runs in the background and serves the model.
  2. Install Open WebUI (in Docker Desktop, or on Linux with the commands below) and open it in a browser. The first account becomes its admin.
  3. In Open WebUI, open Admin Panel → Settings → Models, choose to manage models, and pull qwen2.5:7b.
  4. Start a chat, pick the model, and ask it something from your own work.

Command line

# 1. The model runtime (Linux)
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen2.5:7b

# 2. The chat staff use, on port 3000
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway \
  -e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
  -v open-webui:/app/backend/data --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

# 3. Check the model answers
ollama run qwen2.5:7b "Summarize the purpose of a council work session in two sentences."

The first account created in Open WebUI becomes its admin. Sign in staff with Active Directory: see Open WebUI with Active Directory permissions.

Connecting it to your systems, safely

A chat that can't see your documents is limited. A chat that sees them through one service account sees everything: HR files, finance, the clerk's records, all of it, for every employee. The safe middle is a gateway that gives each person's AI exactly that person's reach.

Acutis Gate does this and installs on the same box. On Windows Server, run the signed MSI, then Start → Acutis Gate → Set up Acutis Gate: it sets up its service account, certificate and browser sign-in policy, and asks nothing unless something needs your decision. On Linux, the package brings Open WebUI and a local model with it (or points at the model you already run), adds HTTPS, and sets up Gate:

Command line

tar xzf AcutisGate-Linux.tar.gz && cd AcutisGate-Linux-*
sudo bash install.sh --name gate.town.example --license ~/acutis-gate.license

From there: staff sign in with their Windows logon, their AI opens file shares as them (only the drives Group Policy maps for them), HR and finance answer by their own permissions, approvals cover changes, and every action is recorded for your SIEM. Then open the Gate console's Health check, which walks the whole chain and names anything left to fix.

Your model, your building, each person's own reach

Gate installs next to your local model: Windows sign-in, files as the person, rules and a full audit trail.

Start a 14-day trial Tour the live Gate

Frequently asked questions

Can a small town run its own AI model?

Yes. A server or mini-PC with 16 GB of RAM runs a 7B open model on CPU, with Open WebUI as the chat. A GPU makes it fast enough for many people at once.

How much RAM does a local LLM need?

About 8 GB for a 3B model, 16 GB for a 7B model and 32 GB for a 14B model, running quantized on CPU.

Is a local LLM safe for resident data?

The model itself keeps data in the building. The risk is what it can reach: connect it to your systems through a gateway that runs each action as the employee who asked, so it never sees more than they could.

Which open model should a school district start with?

A tool-calling model sized to your RAM, such as qwen2.5:7b on 16 GB. Try it on real tasks for a week before buying bigger hardware.