📄
docura powered bydeadbytesIntelligent Document Understanding at Scale
Docura is your AI-powered backend for automating queries, summaries, and insights from documents — built with FastAPI, FAISS, and RAG, and supercharged by deadbytes.
git clone https://github.com/msnabiel/Docura.git
Tech Stack & Features
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
What makes Docura Better?
Smarter retrieval, cleaner output, less guesswork.
State-of-the-Art Extraction
- Hybrid architecture integrating OCR, layout analysis, and NLP-driven content extraction
- Robust multi-engine parsing (OCR, PDF, layout) with intelligent fallback strategies
- Streamlined preprocessing: noise filtering, normalization, and encoding repair
- Context-aware chunking with NLTK, layout preservation, and smart de-hyphenation
- Comprehensive support for DOCX, PPTX, and mixed-format files with table and comment handling
Agentic Intelligence
- Performs intent classification to detect whether the input is a document-related query or an action-triggering command
- Identifies required parameters using NLP parsing and reranks relevant chunks from hybrid vector-keyword retrieval
- Reiterates over retrieved context (e.g., via FAISS) to auto-fill missing or incomplete parameters
- Executes the appropriate backend API call using the resolved intent and structured parameter set
Optimized Hybrid RAG
- Combines smart semantic search with keyword matching to find the most relevant information
- Reranks results to improve answer accuracy using a powerful language model
- Merges and filters content chunks to avoid repetition and keep responses clean
- Keeps track of conversation history for contextual understanding
- Pre-processes documents in advance for faster response times
- Built for speed and scalability, works seamlessly with models like Gemini
Coming Soon to Docura
Discover the next generation of document intelligence — interactive, multilingual, and voice-driven.
Multi-User Chat
Collaborate with teammates in shared AI-powered chat rooms — co-ask, discuss, and refine answers together in real time.
Multimodal Intelligence
Process and understand videos, and audio with advanced embeddings — search and analyze across formats.
Voice Commands
Control your document assistant with natural voice queries — fast, hands-free, and contextual.