A High-Performance, Local Retrieval-Augmented Generation (RAG) Solution with Modern UI.
Optimized for extreme speed and precision usingLangGraphorchestration, localOllamamodels, and a sleek Streamlit interface.
GraphRAG-Ollama features a refined sidebar, real-time status logging, and a professional PDF viewer with a grouped-control navigation toolbar.
- Header-Aware Semantic Chunking: Beyond simple text splitting, our system respects Markdown headers (
#,##). It prevents context contamination between sections and injects structural metadata (current_section) into every chunk. - 99.8% Metadata Optimization: Implements Reference-based Metadata Offloading. Word coordinates for highlighting are stored in a dedicated side-cache (
CoordCacheManager), reducing FAISS index RAM usage by over 99% while maintaining sub-millisecond hydration during retrieval. - FlashRank Semantic Reranking: Integrated
FlashRank(v0.2.0) with ONNX runtime. Re-evaluates search results using a Cross-Encoder model on CPU, achieving 2x higher accuracy (P@1) with minimal latency.
- Configurable Prompt System: All RAG prompts (
grading,rewriting,QA) are externalized inconfig.yml. They use generalized instructions to handle complex model names and technical terms without hardcoding. - Section-Aware Citations: Implements an advanced citation processor that extracts section names and matches them with document metadata. Provides Interactive Citation Badges with sub-second preview tooltips.
- Relevant Entity Extraction: The LLM is strictly instructed to extract key technical entities during the grading phase, ensuring transparency in its reasoning process.
- Multi-Layer Unit Testing: 13+ new unit tests covering Graph Flow, RAG Orchestration, Document Processing, and Advanced Citations.
- 100% Integrity Pass: Integrated verification system (
verify_integrity.py) ensures code style, typing, and RAG logic are always production-ready. - Ghosting-Free UI: Advanced
st.fragmentandst.emptyplaceholder management prevents visual glitches during real-time streaming.
- Streamlit: 1.60.0
- LangChain: 0.3.18
- LangGraph: 0.2.74
- PyMuPDF4LLM: latest
- Ollama: 0.6.1
- FastAPI: 0.133.1
rag-system-ollama/
├── src/
│ ├── .deepeval/
│ ├── api/
│ ├── cache/
│ ├── common/
│ ├── core/
│ ├── data/
│ ├── infra/
│ ├── main.py # 🏁 Entry Point
│ ├── security/
│ ├── services/
│ └── ui/
├── scripts/
│ ├── .model_cache/
│ ├── analyze_dom.py
│ ├── analyze_logs.py
│ ├── analyze_paths.py
│ ├── bench_ui_render.py
│ ├── bench_ui_render_v2.py
│ ├── benchmarks/
│ ├── check_css_presence.py
│ ├── compare_chunking_logic.py
│ ├── container_dom_test.py
│ ├── debug_layout.py
│ ├── deep_dom_analysis.py
│ ├── diagnose_input.py
│ ├── diagnose_ui.py
│ ├── discover_selectors.py
│ ├── dump_dom.py
│ ├── dump_page.py
│ ├── e2e_performance_benchmark.py
│ ├── eval_grader.py
│ ├── eval_results.json
│ ├── eval_retrieval.py
│ ├── evaluate_pipeline.py
│ ├── evaluation/
│ ├── explore_dom.py
│ ├── find_containers.py
│ ├── inspect_containers.py
│ ├── kill_streamlit.py
│ ├── maintenance/
│ ├── quick_verify_rag.py
│ ├── README.md
│ ├── reverify_dom_task1.py
│ ├── simple_eval.py
│ ├── standardize_imports.py
│ ├── test_app.py
│ ├── test_embedding_v2.py
│ ├── test_full_pipeline.py
│ ├── test_highlight_query_cleaning.py
│ ├── test_pipeline_direct.py
│ ├── test_rag_eval.py
│ ├── test_real_pdf_chunking.py
│ ├── test_self_correction.py
│ ├── validate_config.py
│ ├── verification/
│ ├── verify_css_override.py
│ ├── verify_dom_structure.py
│ ├── verify_e2e_all.py
│ ├── verify_final.py
│ ├── verify_fixes.py
│ ├── verify_height_fill.py
│ ├── verify_layout_height_fill.py
│ ├── verify_layout_scroll_fix.py
│ ├── verify_metadata_opt.py
│ ├── verify_new_layout.py
│ ├── verify_phase2.py
│ ├── verify_section_metadata.py
│ ├── verify_styles.py
│ ├── verify_ui_scrolling.py
│ └── visual_qa_automation.py
├── tests/
│ ├── conftest.py
│ ├── data/
│ ├── e2e/
│ ├── integration/
│ ├── README.md
│ ├── security/
│ ├── smoke_test.py
│ ├── stability/
│ ├── test_loop_independence.py
│ ├── test_p0_1_pdf_handle_leak.py
│ ├── test_ui_bridge.py
│ ├── unit/
│ └── utils/
# Pull the recommended models
ollama pull qwen3:4b-instruct-2507-q4_K_M
ollama pull nomic-embed-text-v2-moe # default embedding model (config.yml: models.default_embedding)
# Optional (legacy embedding model):
ollama pull nomic-embed-textCustomize your RAG behavior in config.yml (e.g., grading instructions, chunk sizes, model names).
# Optimized for Windows environments
streamlit run src/main.pyThe API server bootstraps an admin account: set TEST_ADMIN_PASSWORD (else a random password is printed to stderr once). All API routes require a Bearer token (JWT via /api/v1/login or an API key). Revocation persists via AUTH_STATE_FILE; tokens survive restarts. API uploads are stored in data/temp/pdf_library and served at /api/v1/pdf/{hash}.
We maintain a strict Zero-Error Policy. Run the automated verification suite:
# Run all unit tests
pytest tests/unit
# Run full pipeline integration test
python scripts/test_full_pipeline.py
# Verify section metadata extraction
python scripts/verify_section_metadata.pyCI enforces unit coverage ≥55% (--cov-fail-under=55) and additionally runs the auth/ownership/PDF/SSE integration tests (test_api_auth_login, test_api_pdf_serving, test_global_exception_handler, test_ownership_hardening, test_pdf_library_retention, test_stream_error_isolation, test_api_endpoints).
Evaluate retrieval and answer quality end to end with the built-in harness:
# Full run: retrieval + generation + judge scoring
python scripts/eval_quality.py --tag <tag> --testset_n 3
# Retrieval-only: skip the LLM judge for faster iteration
python scripts/eval_quality.py --tag <tag> --no-llmEach run writes reports/eval_quality_<tag>_<ts>.json and .md with per-question metrics (P@1, MRR@5, TTFT, tokens/sec) plus an average judge score.
Reranking. The reranker engine is selected in config.yml → rag.reranker.engine: auto (tries the FlashRank cross-encoder, falls back to the bi-encoder semantic reranker on failure), flashrank (forced), or semantic (bi-encoder). When a pipeline is built, the default LLM is preloaded automatically (silently skipped if loading fails).
Key settings:
| Key | Default | Purpose |
|---|---|---|
models.ollama_num_predict |
2048 | Max generated tokens per answer |
models.num_ctx |
8192 | Model context window (tokens) |
rag.prompts.grading.min_score_to_skip |
0.85 | Skip the LLM grade when the max rerank score reaches this |
ui.timeline_poll_seconds |
1.0 | Timeline fragment auto-refresh interval (seconds) |
MIT License - Developed by darkzard05. Status: v3.3.0 | Last Updated: 2026-03-12
