idun-gguf β Local GGUF Tool-Agent
idun-gguf is a full functional local Tool-Agent, that mirrors the pattern of idun-sdk 1:1 β with any local GGUF-Model via Ollama instead of Azure Cloud.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
β idun-gguf CLI / Python / MCP β
β (OpenAI-compatible Client, stdlib-only) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Ollama (localhost:11434) β
β qwen3:8b (GGUF, Q4_K_M) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Tool Registry (7 Tools) β
β web_search memory file_ops code_executor β
ββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Full Agent Trajectory Output β
β .text (final Answer) + .steps β
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Requirements
- Python β₯ 3.10
- Ollama
- ~8GB free RAM (qwen3:8b)
- ~5.2GB Diskspace (qwen3:8b)
Installation
1. Ollama
curl -fsSL https://ollama.com/install.sh | sh
2. Ollama Server
ollama serve &
3. Pull the Model
ollama pull qwen3:8b
Notice:
qwen3:8bhas excellent Tool-Calling (5.2GB, Q4_K_M).
Alternatives:qwen3:14b,llama3.1:8b,mistral-nemo:12b
4. Install idun-gguf
# From the project directory
cd /app/idun_gguf_integration_1709
pip install -e .
# Direct install
pip install -e /app/idun_gguf_integration_1709/
5. Dependencies
pip install duckduckgo-search chromadb click
Quick Start
CLI
# Simple Question (final Answer)
idun-gguf chat "What is Python?"
# Full Agent-Trajectory
idun-gguf trace "What is Python?"
# Show available tools
idun-gguf tools
# Show defined model
idun-gguf model
Python Client
from idun_gguf import IdunLocalClient, Response
# Client create
client = IdunLocalClient()
# Simple answer
response = client.complete("What is Python?")
print(response.text) # Final Answer
print(len(response.steps)) # Step count
# With Trajectory
for step in response.steps:
print(f"[{step.type}] {step.content}")
if step.type == "tool_call":
print(f" β Tool: {step.tool_name}, Input: {step.tool_input}")
if step.type == "tool_result":
print(f" β Ergebnis: {step.tool_output[:100]}...")
MCP Server
# Tools listing
echo '{"jsonrpc":"2.0","method":"tools/list","id":1}' | \
python3 -m idun_gguf.mcp_server
# idun_chat call
echo '{"jsonrpc":"2.0","method":"tools/call","params":{"name":"idun_chat","arguments":{"prompt":"Was ist Python?"}},"id":2}' | \
python3 -m idun_gguf.mcp_server
CLI Commands
idun-gguf chat
# with prompt as Argument
idun-gguf chat "What is Python?"
# with System-Prompt
idun-gguf chat "Describe ML" --system "You are an ML-Expert"
# as JSON-Output
idun-gguf chat "Hello" --json
# with another Model
idun-gguf chat "Hello" --model llama3.1:8b
idun-gguf trace
# Display complete Trace
idun-gguf trace "What is Python?"
# As JSON-Output
idun-gguf trace "Hello" --json
idun-gguf tools
idun-gguf tools
idun-gguf modetime
idun-gguf model
Tools
| Tool | Description | Parameters |
|---|---|---|
web_search |
DuckDuckGo-Websearch | query, max_results |
memory_search |
Vector-Memory indexing | query, n_results |
memory_save |
Save info in Vector-Memory | text, source |
file_read |
Data read | path, max_lines |
file_write |
Data write | path, content |
file_list |
List directory | path, pattern |
code_executor |
Python-Code execute | code, timeout |
Architecture
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β idun-gguf Package β
β β
β βββββββββββββββ βββββββββββββββββββ β
β β cli.py β β mcp_server.py β β
β β (Click) β β (JSON-RPC) β β
β ββββββββ¬βββββββ ββββββββββ¬βββββββββ β
β β β β
β ββββββββ΄βββββββββββββββββββ΄βββββββββ β
β β client.py β β
β β IdunLocalClient: .complete() β β
β β β Response(.text, .steps) β β
β βββββββββββββββββ¬βββββββββββββββββββ β
β β β
β βββββββββββββββββ΄βββββββββββββββββββ β
β β Tool Registry β β
β β web_search β memory β file_ops β β
β β code_executor β β
β ββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββ
β Ollama (localhost) β
β qwen3:8b (GGUF) β
β OpenAI-compatible API β
βββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββ
β GGUF Model (Q4_K_M) β
β ~5.2GB, CPU-Inference β
βββββββββββββββββββββββββ
Data flow
- User Prompt input (CLI / Python / MCP)
- IdunLocalClient Send Prompt + Tool-Definitionen to Ollama
- Ollama Running the Model (qwen3:8b)
- Model decides: Answering or Tool-Call
- Tool Registry Execute tool-calls from (web_search, memory, etc.)
- Results are passed back to the model
- Loop repeats until final answer
- Response with
.text(final answer) +.steps(complete trajectory)
Environment variables
| Variable | Default | Description |
|---|---|---|
OLLAMA_BASE_URL |
http://localhost:11434 |
Ollama Server-URL |
IDUN_GGUF_MODEL |
qwen3:8b |
Standard-Model |
Compare
| Aspect | idun-sdk (Azure) | idun-gguf (Local) |
|---|---|---|
| Backend | Azure AI Foundry | Ollama + GGUF |
| Model | NatureLM-Idun-5-MoE | qwen3:8b (changeable) |
| Latency | 2-10s (Cloud) | 30-120s (CPU) |
| Cost | Pay-per-Token | Free (Electricity) |
| Data Security | Cloud | 100% Local |
| Offline | β | β |
| Tools | web_search, memory_search | 7 Tools (expandable) |
| Trajectory | β .text + .steps | β .text + .steps |
Troubleshooting
Ollama check
# Check
curl http://localhost:11434/api/tags
# Restart
ollama serve &
Model not found
ollama list # Check if qwen3:8b is listed
ollama pull qwen3:8b # if not listed then pull qwen3:8b
Package-Import-Failure
pip install -e /app/idun_gguf_integration_1709/ --force-reinstall
ChromaDB-Failure
# ChromaDB persistent in ~/.idun_gguf_memory/
# Bei Problemen: rm -rf ~/.idun_gguf_memory/
Development
# Repository clone
git clone https://huggingface.co/Qapdex/Idun-5-MoE-sdk-GGUF
cd idun-gguf
# Developermode
pip install -e .
# Tests
python -m pytest tests/
License
MIT
Related projects
- idun-sdk β Thin Client for NatureLM-Idun-5-MoE on Azure