Retrieval-Augmented Generation (RAG) systems address the hallucination problem in large language models by grounding responses in retrieved passages from trusted sources, using hybrid retrieval combining dense vector search with sparse lexical search (BM25), cross-encoder reranking for precise ordering, and question decomposition agents to handle complex multi-part queries, while implementing faithfulness verification to ensure answers are supported by retrieved evidence.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
video1733712554
Added:Hello sir myself and I am from IIT trimester. Currently I am presenting my trimester 7 project uh which I I have done. So my project name was Li Lia a fully local aentic retrieval augmented generation system. Uh so my name is Praarai and my student ID is 2305010.
So what the problem and what the motivation uh of doing this project was?
Uh so the problem is large language models are fluent but not grounded ungrounded and uh if we asked about a document they have never seen the illusion and possible falsehoods. So what happen is sometime if we upload a PDF and then we ask something related to the PDF or any financial PDF or any PDF then the system uh searches and answer from it but if we not seen any PDF and we have not uploaded any PDF and if you ask about that PDF then also it answers something so it so the retrieval of generation fixes this fly rounding answers in the retrieved passages from the trusted source. But the name receip embed take top k by sign stop the prompt break on real document. So what the problem it solved is the rag system.
The rack system it uh embeds and it takes stock case for sign similarity stuff the prompts and break break on real documents such as vocabulary mismatch dense vector blur exact terms names and ids for course ordering independent similarity scores misrank by sellers and single sort irrial one query cannot cover multi- part question so the goal of the our project is to accurate retrievable and multi of reasoning and verifiable routing.
So the objective of the project is build an end toend direct product over user PDFs, URLs and notes and answering with inline citations. It implement hybrid retrieval dense eventing sparse lexical used without consiling score scales.
Add cross encoder reranking for precise ordering of the sort list. Then introduce an agent that decomposes complex questions into multiop subqueries.
Then it guard against vucination with an automated pathfulness check and keep it fully local open and compared uh 340 lines with the fair tier zone. So we will use T4 GPU in this project. So what is the architecture? Simple system architecture of his project is the user feeds it question. Uh then uh the rank plan of queries then dance plus sparse retrie uh RF plus then synthesize cited answer and faithfulness check and after that plan retrie and rerank synthesize and then verify the answer and then it gives the result in JSON and indexing structure our chunking pack the paragraphs into two 20 word windows with a 40 word sliding Overlap oversized paragraph are hard split so nothing overflows. Then each chunk keeps its source so answers cannot site it.
Dense index busy small embedding to normalize search by fast in the inner product and captures meaning. A sparse index that OKAP BM25 over the same chunks capture exact terms. two complimentary use of the corpus build once at in then we move on the next step of this project which is hybrid retrieval plus reciprocal rankus through the dense generalizes across pars exact terms their failure modes are complimentary then it merge the 2D rank list with reciprocal rank fusion using only one rank position is 4D D is equal to R which belongs to the set of R one up for K + rank R of D and K is equal to 60 because RRF ignores raw scores it slide steps reconciling cosine similarity with BM25 magnitudes robust and parameter light re-ranking and aentic reasoning cross encoder reanking report the fuel sort list by reading query plus passage together far more accurate than independent by encoder scores applied only to sort list. So the accurate model stays affordable.
Then aantic planning the LLM decomposes a complex question into three subqueries for already simple and retrieve for hope then pull and dduplicate the evidence.
This targets exactly where nag is requested compositional questions.
Then we move to the next step which is grounding and faithfulness verification.
So the model answer synthesis statement.
The model answer using only the number sources and attaches an inline citation end to every claim verification and facts checking pass compare the answer against the sources and returns online.
Okay. Grounded a partially unsupported naming any unsupported claim. When evidence is left in the model decline rather than fabricates and result is a visible actionable trust signal on every onsite implementation and result stag 2.53 helium instruct 4 n4 bg small mini L cross encoder fires BM25 gradu all open and t4. So these were all the stacks of what the technology we are using. So we use the model Q23 instruct to 4bit NF4 uh BGE small mini cross and folder fires as our open fast API DM25 and gradio for implement deploying our opensource app. So the hybrid RF what it recovers both exact terms and regress plus reanking most relevant passage rises to the top planning better coverage uh of multi-art questions verification it lacks unsupported cla to the user and put three lines run fully the largest gains come from retable and rounding engineering not from global model size.
So what conclusion and future work we can do on this project is a generally useful production saved R product built compar compactly on free open rate infrastructure with answers traceable to users own sources. Uh the future work is conversational memory for followup questions quantitative evaluation arness retrie and performance iterative self correction when the ffulness check fails and persistent with incremental vector for large part.
Thank you.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Bitcoin Social Interest: Dozens of us Left
benjaminjcowen
12K views•2026-07-23

Tesla Profits Plunge & SpaceX Stock Continues Fall
TheJohnJohnstonLounge
6K views•2026-07-23