Rabu, 30 September 2026

Local AI Weekly #4: The Fine Print of Running AI Locally

local ai weekly

Welcome to Local AI Weekly #4.

"Local AI" keeps meaning different things these days. Sometimes the model runs on your computer. Sometimes only the app does, while the model and your chat still run elsewhere.

There are lots of such "local AI" tools that are not really local. Before you pick a local AI tool, ask three things: where does inference happen, what account or network service is still required, and what does the license let you do? The best local AI tool is the one that doesn't need any of that.

🧪 On my bench

Team It's FOSS has moved from Discord to Buzz for internal chat. Buzz, from Jack Dorsey of Twitter fame, is a decentralized communication tool for humans and agents. Agents can run on remote servers or a local harness over ACP.

It has quirks. Clipboard screenshots won't paste into chat, and desktop notifications only fire for direct messages. Manageable for now.

🔍 Discover AI tools

Kubutu developer Rick Timmis is working on Klara, a local desktop AI assistant for KDE Plasma. The idea is to let you control the desktop via voice input. It is a work in progress for now.

OpenMuse is an MIT-licensed personal-agent app with a browser worker, durable tasks, and an optional Docker-based Linux computer. You can host it yourself, but setup still needs a CopilotKit Intelligence key, and open-ended tasks default to cloud providers. You can point it at an OpenAI-compatible endpoint, but there's no documented native Ollama path. Self-hosted software, not an assured offline agent.

🎫 Get MCP Certified

The Linux Foundation now offers AI certifications. The Model Context Protocol Associate one could interest you if you're building MCP integrations, or want an AI credential on your resume. I plan to take it just for the sake of learning new skills.

📡 Open Model News

OpenDecider is a more modest use for local models: answer bounded questions, like which queue a ticket should go to, instead of writing long replies. Its nano model is about 400M parameters; the 4B small model is a Qwen-based adapter. The author says both were distilled from larger teachers, and reports 2.0 GiB for nano and 8.9 GiB for small in the tested setups.

Basically, you don't always need a giant model to route routine requests. A small student model can be the triage step ahead of a slower agent.

👀 Big Tech Watch

NVIDIA's September PAIR announcement also promises easier local-model setup in Hermes and OpenClaw. The one-click Hermes path launched on Windows, with Linux "coming soon." Don't confuse that future path with the Linux PAIR beta available now.

NVIDIA also quotes up to 1.9x throughput from llama.cpp optimizations on an RTX 5090. That's a vendor number on specific hardware, not a speedup I'd expect on yours.

🗂 AI Jargon: Distillation

Imagine you have a giant, super-smart teacher who knows everything about the world. This teacher has a massive brain, but its so big that it can only stay inside a giant school building.

AI distillation is like that big teacher sharing all their secrets with a little kid (the student).

Instead of making the kid read millions of textbooks, the big teacher says: "Don't worry, just watch how I solve these puzzles, and listen to how I think."

The little kid watches closely and learns the teacher's smart shortcuts. Soon, the kid becomes almost as smart as the teacher, but with a much smaller brain!

You can learn more about distillation here. And you will see that teacher-student is kind of official term in this context.

😂 Meme

When the open-weights drop looks a little too familiar...

AI meme

⚡ Quick Tip: Back up your Hermes agent before you need to rebuild the harness

Last issue: ollama ps, to see if your model was really on the GPU. This week, make sure you could rebuild the setup around it too.

If you run Hermes, run hermes backup. It writes a ZIP of your Hermes home, config and state included, restored later with hermes import path/to/backup.zip. hermes backup --keep N caps retained backups, and a script-only cron job runs it on schedule without starting an agent.

Then move one encrypted copy off the machine. That archive can hold credentials, sessions, memory, and config, so don't drop the raw ZIP in Git, even a private repo. It won't be wise.

If you have not subscribed to Local AI Weekly yet, you can subscribe from this page.

Subscribe to Local AI Weekly

See you next week.



from It's FOSS https://ift.tt/2zLDaTK
via IFTTT

Tidak ada komentar:

Posting Komentar