Running large language models locally with Ollama — VRAM requirements by tier, quantization formats explained, model selection for different use cases, Open WebUI setup, modelfiles for custom system prompts, and an honest comparison of when local beats cloud and when it doesn't.
Local-Ai
-
Local LLMs with Ollama -
Local LLM Deep Dive: Ollama, Quantization, and Running AI on Your Own Hardware A comprehensive technical guide to running large language models on your own hardware — covering Ollama setup, quantization formats, hardware selection, Apple Silicon and NVIDIA GPU configuration, and building a private local coding assistant.
-
RAG from Scratch: Building a Retrieval-Augmented Chatbot Over Your Own Documents A complete technical walkthrough of building a retrieval-augmented generation (RAG) system from scratch — covering embeddings, vector databases, chunking strategy, retrieval quality, and a fully runnable Python implementation using Ollama and ChromaDB.
-
Local LLM Inference for Coding: The Complete 2025/2026 Guide Everything you need to run powerful AI coding assistants locally — model benchmarks, Ollama, LM Studio, llama.cpp, hardware requirements, and editor integration with Continue.dev and Aider.