A fully local RAG (Retrieval Augmented Generation) CLI application built in Go. Ingest markdown files, create embeddings, store them in Milvus, and chat with your documents using Ollama.
docker compose up -dWait for Milvus to be healthy:
docker compose psollama pull nomic-embed-text
ollama pull llama3.2go build -o docmind ../docmind ingest ./docsThis will:
- Scan the directory for
.mdfiles - Split them into heading-aware chunks
- Generate embeddings via Ollama
- Store everything in Milvus
./docmind chatType your questions and get answers grounded in your ingested documents. Type quit to exit.
All settings can be overridden with environment variables:
| Variable | Default | Description |
|---|---|---|
OLLAMA_URL |
http://localhost:11434 |
Ollama API URL |
MILVUS_ADDR |
localhost:19530 |
Milvus gRPC address |
EMBED_MODEL |
nomic-embed-text |
Ollama embedding model |
CHAT_MODEL |
llama3.2 |
Ollama chat model |
CHUNK_SIZE |
512 |
Max chunk size in characters |
TOP_K |
5 |
Number of results to retrieve |
Ingest: .md files -> Chunker -> Ollama Embedder -> Milvus
Chat: Question -> Embed -> Milvus Search -> Build Prompt -> Ollama LLM -> Stream Response