← All guides

Guides · Local AI

Local AI: what it is, what it needs, and when it's worth it.

Running AI on a computer you own sounds technical, and parts of it are. The decision behind it is simpler. Here's what local AI can do for a small business, what it takes, and when a cloud tool is the better call.

An accountant has a client's email thread open: forty messages about a messy year, three revised numbers, and an attachment nobody can find. An AI tool could summarize it in about ten seconds. She isn't comfortable pasting a client's financial details into a public chatbot, so she reads all forty messages instead.

That hesitation is reasonable, and it's common. Bookkeepers, clinics, law offices, and anyone holding other people's information feel the same pull. The tools are useful, and the data isn't theirs to hand around. Local AI is one way through. It has real tradeoffs, and it isn't the right answer for everyone, but it's worth understanding before deciding either way.

What "local AI" means

Local AI means the AI model runs on a computer you own, in your office or at home, instead of on a company's servers. When you ask it to summarize a document, the document and the question are processed on that machine. Set up properly, nothing about the request goes to an outside AI service.

The cloud tools most people know, like ChatGPT, Claude, and Gemini, work differently. Your request travels over the internet to the provider, gets processed there, and the answer comes back. Those tools are fast and capable, and business plans come with meaningful privacy commitments. The processing still happens somewhere else, under someone else's terms.

Local setups usually run open-weight models: versions that companies such as Meta, Google, Mistral, and Alibaba have released for anyone to download and run. Free programs like Ollama and LM Studio make running them much simpler than it was a few years ago. LM Studio, for example, dropped its separate commercial licence in 2025 and is now free to use at work.

Why a small business would want it

Privacy and control

This is the main reason. Client files, health information, financial records, and contracts stay on hardware you control. For a business with confidentiality duties, that can make AI usable where it otherwise wouldn't be.

Predictable cost

Once the hardware is paid for, there's no per-seat subscription and no usage bill. The cost moves from a monthly line item to a one-time purchase, plus electricity and a little upkeep.

No dependence on someone else's roadmap

A local setup keeps working the way you configured it. It doesn't change because a provider adjusted its pricing, retired a model, or updated its terms. It also runs without an internet connection, which matters less than it sounds for most offices, but it's a nice property to have.

The useful question isn't whether local AI beats the cloud. It's whether the information you'd use it on should leave the building.

What it does well today

Local models are strongest at work that's mostly reading and rearranging text you already have:

Those jobs share a pattern. The information is already in front of the model, and the task is to organize it. That's where smaller models hold up well.

How "ask our documents" works

The most requested job is some version of "let me ask questions about our files." The usual approach rests on a plain idea. Your documents are split into small passages and indexed. When someone asks a question, the system finds the passages most likely to hold the answer and gives only those to the model, along with the question. The model answers from what it was handed, and a good setup shows which document the answer came from.

Two practical consequences follow. Answers are only as current as the documents you've added, so someone needs to keep the folder up to date. And messy source material makes for messy answers: a policy manual with three conflicting versions will produce conflicting answers until the duplicates are cleaned up.

Where it falls short

This is the part that sales pages tend to skip.

Smaller models are less capable

The models that run comfortably on one office computer are smaller than the ones behind the leading cloud services. They can be noticeably weaker at complex reasoning, nuanced writing, and long, multi-step tasks. Summarizing a contract is a reasonable job for them. Drafting the argument in one still needs a person doing the thinking.

Speed depends on the hardware

On modest hardware, answers arrive slowly, a few words at a time. On the right hardware, they're quick enough that nobody notices. The difference is mostly memory, which is the next section.

Someone has to own it

A local setup needs updates, backups, and occasional attention. It's a small job, but it's a real one, and it needs a name next to it.

It still makes mistakes

Local models can be confidently wrong, the same as cloud models. Anything going to a client should be read by a person first.

What hardware it needs

The most important number is memory. A model has to fit into memory to run well, and on most PCs that means the graphics card's memory, called VRAM. On Apple silicon Macs, the processor and graphics share one pool of unified memory, which is why a Mac with a lot of memory can run surprisingly large models.

Model size is measured in parameters, usually in the billions. Most local setups use quantized versions of models, compressed enough to fit on everyday hardware with a modest loss in quality. As a rough guide, here's what a few sizes take, using one popular model family as the example:

Model sizeDownload sizeA comfortable starting point
4 billion parameters2.5 GBMost recent computers; fine for testing
8 billion5.2 GBA graphics card with 8 GB of VRAM
14 billion9.3 GBA graphics card with 12 to 16 GB
32 billion20 GBA graphics card with 24 GB, or a Mac with 48 GB or more

Download sizes are Ollama's listings for the Qwen3 family, checked in October 2026. Treat the right-hand column as a starting point. Longer documents need extra memory on top of the model itself, and other model families vary in size.

In practice, a capable small model runs on a mid-range graphics card. The larger models that feel closer to cloud quality need high-end cards or a lot of unified memory, and that's where the budget climbs. The rest of the machine matters less, though 32 GB of system memory and a fast SSD with room for several multi-gigabyte models will save you headaches.

A sensible first setup for an office

For most small offices, it looks something like this:

  1. One capable desktop, kept in the office, with a graphics card sized to the models you plan to use.
  2. A free model runner such as Ollama or LM Studio, with one or two models chosen for the tasks you have in mind.
  3. A simple chat window that staff reach over the office network, so nobody installs anything on their own computer.
  4. A short written rule covering what goes in, who maintains it, and how it's backed up.

Before buying anything, test on a computer you already have. A small model on an existing laptop will tell you a lot about whether the idea fits your work, even if it runs slowly.

Two things you can try yourself, this week

Neither one needs a purchase, and both will make the decision clearer.

The prep

Write your "never paste" list. Spend ten minutes listing the documents and details you wouldn't put into an online AI tool: client financials, health information, contracts, employee records. That list does two jobs. It's the start of a simple AI policy for your team, and it's the clearest picture of where local AI would earn its keep.

The fix

Try a small model on what you already own. Install LM Studio or Ollama, download a small model, and give it a document with nothing sensitive in it, like a public policy or an old newsletter. Ask for a summary and a list of action items. If it's slow, you've learned which hardware matters. If it's fine, you may not need new hardware yet.

Keeping a local setup secure

Local doesn't mean secure by default. A few basics do most of the work:

A note on Canadian privacy rules

Running AI locally can make privacy simpler. It doesn't remove your obligations. In Ontario, private-sector businesses that handle personal information in the course of commercial activity fall under PIPEDA, the federal privacy law. Under PIPEDA, an organization stays accountable for personal information it hands to a third party for processing, and that includes a cloud service. Keeping the processing on your own machine means fewer parties to account for.

The rest still applies: consent for how you use information, reasonable security, and sensible retention. Health information custodians in Ontario, such as dentists, physiotherapists, and chiropractors, have further duties under the Personal Health Information Protection Act (PHIPA). If your business handles sensitive information, talk to a privacy professional about your situation. This guide isn't legal advice.

Is local AI right for you?

It's probably a good fit if:

A cloud tool is probably the better call if:

Plenty of businesses end up with both: a cloud tool for general work, and a local setup for the files that shouldn't leave.

Common questions

Is local AI completely private?

It's as private as the machine and network it runs on. Physical security, updates, passwords, and backups still matter. Local processing removes the outside AI provider from the picture, but a poorly secured computer is still a poorly secured computer.

Do I need a gaming PC to run AI locally?

You need a lot of graphics memory, and gaming graphics cards are often the most affordable way to get it. The rest of the machine can be ordinary. A Mac with plenty of unified memory is another route.

Can a local model match ChatGPT or Claude?

Not the largest versions, on typical office hardware. For focused jobs like summarizing, drafting from templates, or answering questions about your own documents, a well-chosen local model can be good enough, and that's the bar that matters.

Does local AI work without an internet connection?

Yes. Once the software and models are downloaded, everything runs offline. You'll want an occasional connection for updates.

Where to start

Write the "never paste" list, then try a small model on a computer you already have. If the idea fits and the hardware is the bottleneck, that's when a dedicated machine makes sense. If you're in Toronto and want help planning one, from choosing the parts to installing the models and connecting them to your documents, that's the kind of build Happy Genius does. For examples of what this looks like in a clinic or office, see AI systems for clinics and professional offices.

Want local AI set up for your office?

Free 30-minute consultation. We'll look at the work you'd use it for, tell you honestly whether local or cloud fits better, and what the hardware would take.

Book My Free Consultation