📄 docura powered bydeadbytes

Intelligent Document Understanding at Scale

Docura is your AI-powered backend for automating queries, summaries, and insights from documents — built with FastAPI, FAISS, and RAG, and supercharged by deadbytes.

git clone https://github.com/msnabiel/Docura.git

Tech Stack & Features

FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments
FAISS
Gemini 2.5 Flash
LangChain
FastAPI
Hugging Face
GGML Inference
NLTK
Smart Chunking
Offline Embeddings
Docx/PDF/Email
Local Vector Store
Terminal-based Deploy
Docker Deployments

What makes Docura Better?
Smarter retrieval, cleaner output, less guesswork.

State-of-the-Art Extraction

  • Hybrid architecture integrating OCR, layout analysis, and NLP-driven content extraction
  • Robust multi-engine parsing (OCR, PDF, layout) with intelligent fallback strategies
  • Streamlined preprocessing: noise filtering, normalization, and encoding repair
  • Context-aware chunking with NLTK, layout preservation, and smart de-hyphenation
  • Comprehensive support for DOCX, PPTX, and mixed-format files with table and comment handling

Agentic Intelligence

  • Performs intent classification to detect whether the input is a document-related query or an action-triggering command
  • Identifies required parameters using NLP parsing and reranks relevant chunks from hybrid vector-keyword retrieval
  • Reiterates over retrieved context (e.g., via FAISS) to auto-fill missing or incomplete parameters
  • Executes the appropriate backend API call using the resolved intent and structured parameter set

Optimized Hybrid RAG

  • Combines smart semantic search with keyword matching to find the most relevant information
  • Reranks results to improve answer accuracy using a powerful language model
  • Merges and filters content chunks to avoid repetition and keep responses clean
  • Keeps track of conversation history for contextual understanding
  • Pre-processes documents in advance for faster response times
  • Built for speed and scalability, works seamlessly with models like Gemini

Coming Soon to Docura

Discover the next generation of document intelligence — interactive, multilingual, and voice-driven.

Multi-User Chat

Collaborate with teammates in shared AI-powered chat rooms — co-ask, discuss, and refine answers together in real time.

Multimodal Intelligence

Process and understand videos, and audio with advanced embeddings — search and analyze across formats.

Voice Commands

Control your document assistant with natural voice queries — fast, hands-free, and contextual.

Ready to transform your documents?

Start using Docura today.

Build your own smart document chatbot with Docura. Upload files, ask questions, and get instant answers — all powered by local embeddings and retrieval-augmented generation.