Published
Key points
- Retrieval-augmented generation (RAG) searches your documents for passages relevant to each question and has a large language model (LLM) answer from them. Proposed by Lewis et al. in 2020, it updates what the system knows by changing the index, without retraining the model.
- RAG accuracy depends first on which documents are indexed and how they are retrieved, not on the model. Japan's IPA guideline names noisy, duplicated and outdated information as causes of wrong answers.
- RAG reduces hallucination but does not eliminate it. Measure retrieval (did it find the right passages?) and generation (is the answer faithful to them, and does it answer the question?) separately, on an evaluation set built from real questions, after every change.
- Access control is the central security issue. The OWASP Top 10 for LLM Applications 2026 calls for document- and chunk-level authorisation inside the index query, before retrieval, and warns of indirect prompt injection through ingested documents.
- Deployment options are a cloud API, a cloud model over private connectivity, or an on-premises (local) LLM, chosen by data sensitivity, volume and in-house capability. Japanese options for local use include NTT's tsuzumi 2, designed to run on a single GPU, and the open models published by NII's LLM-jp.
What is RAG (retrieval-augmented generation)?
Retrieval-augmented generation (RAG) is a method in which, for every question, a system searches documents or databases for relevant passages and has a large language model (LLM) generate its answer from them. Because the answer rests on material the model was never trained on, such as internal policies, product manuals, contracts and support records, RAG is the typical architecture for generative AI that works over company documents.
The method comes from a 2020 paper by Lewis et al., which proposed RAG as a language generation model combining pre-trained parametric memory (the language model) with non-parametric memory (a searchable index of documents). The appendix to Japan's AI Guidelines for Business (Ver1.2), issued by the Ministry of Internal Affairs and Communications (MIC) and the Ministry of Economy, Trade and Industry (METI), cites this definition and notes that companies use RAG to search internal documents and databases to make generative AI answers more accurate.
The same appendix lists the expected benefits: fewer hallucinations, more transparent answers because sources can be shown, and lower cost because data sources can be added without retraining. The uses cut across industries:
- Manufacturing: answering from maintenance manuals and fault records, citing the procedure and page
- Legal and procurement: finding and quoting the relevant clause in contracts and terms of trade
- Customer support: drafting answers from product specifications and FAQs, with sources
- Internal help desk: answering questions on work rules, expense policy and information security policy
- Financial services: answering enquiries across operating procedures, product documents and internal rules
RAG can only answer from what is in the index. The guideline on introducing and operating text-generation AI from Japan's Information-technology Promotion Agency (IPA) makes the same point: RAG improves accuracy only for information stored in the vector database, so users need to be told what the system can answer.
How RAG works: ingestion, retrieval and generation
A RAG system runs in three stages: ingestion prepares documents for search, retrieval finds the passages relevant to a question, and generation writes an answer from them. Errors surface in the generated answer, but their cause often lies in the first two stages.
- Ingestion (indexing): extract text from PDFs, Word files, wikis and other sources, split it into short units called chunks, convert each chunk into a vector with an embedding model, and store it in a vector database or search engine with metadata such as document name, version and access permissions.
- Retrieval: convert the question into a vector with the same embedding model and fetch the few chunks closest in meaning. Many systems add keyword search, or re-rank the candidates before passing them on.
- Generation: pass the retrieved chunks to the LLM with the question, instructing it to answer only from the material provided and to cite its sources.
In Lewis et al.'s experiments, Wikipedia was split into disjoint 100-word chunks, giving an index of 21 million passages. Holding the index vectors took about 100 GB of CPU memory, reduced to 36 GB with compression. An on-premises build therefore has to budget memory and storage for the index as well as GPUs for the model.
The same paper swapped a December 2016 Wikipedia index for a December 2018 one and asked about 82 world leaders who had changed in between. With the index matching the year, RAG answered 70% (2016) and 68% (2018) correctly; with the indices mismatched, accuracy fell to 12% and 4%. Updating knowledge by replacing the index, without retraining, is a strong reason to choose RAG where policies and product information change often.
RAG vs fine-tuning: which to use when
If you want answers grounded in your organisation's knowledge, start with RAG. Fine-tuning adjusts a model's weights through further training. It suits adapting style, output format or a specific task, but updating knowledge means retraining, and it is hard to show which document an answer came from. Lewis et al. describe both provenance and knowledge updates as open problems for models that hold knowledge only in their parameters.
| Aspect | RAG | Fine-tuning |
|---|---|---|
| Updating knowledge | Change the indexed data | Retrain the model |
| Showing sources | Can cite the documents and passages used | Hard to trace an answer to a document |
| Best suited to | Answering from changing knowledge such as policies, manuals and contracts | Consistent style or format, classification and other specific tasks |
| Main preparation | Cleaning and chunking documents, attaching permissions | Building and quality-checking training data |
| Main risks | Missed retrieval, stale documents, access beyond permissions | Memorised training data reappearing in outputs, retraining effort |
IPA's guideline reports the view that, as of June 2024, RAG was also easier to introduce and more realistic for organisations to use. The two are not mutually exclusive: RAG can supply the knowledge while fine-tuning shapes the output format. The Generative AI Profile from the US National Institute of Standards and Technology (NIST AI 600-1) asks organisations to re-assess model risks after implementing either.
Designing for accuracy: documents, chunking, retrieval and instructions
RAG accuracy is set by the quality of the indexed documents and the retrieval design before the model's capability comes into play. IPA's guideline warns that noisy, duplicated or outdated information in the vector database leads to hallucinations and wrong answers. The appendix to the AI Guidelines for Business adds that RAG accuracy depends on the information it uses, so sources should be reliable and their quality managed and monitored regularly.
Prepare the documents
- Narrow the scope: start with one business process and one document set, and establish where the authoritative current version lives
- Remove old versions and duplicates: keep superseded policies and copies out of the index, and record version and effective date as metadata
- Check extraction quality: sample scanned PDFs, tables and multi-column layouts to confirm the text came out intact
- Attach permissions: record, per document and where needed per chunk, which departments and roles may see it
Design chunking and retrieval
Split chunks at the document's natural boundaries, such as headings, clauses and procedure steps, and attach the document title and heading to each chunk so it still makes sense on its own. Japanese text has no spaces between words, so for Japanese documents measure chunk length in characters or tokens rather than words. A small overlap between neighbouring chunks helps avoid losing information at the boundaries.
Do not rely on vector search alone: consider combining it with keyword search such as BM25. In Lewis et al.'s experiments, BM25 retrieval gave the best result on FEVER, an entity-centric fact-verification task. Questions containing model numbers, clause numbers, product names or customer names benefit from exact keyword matches.
The number of chunks passed to the LLM is another setting to tune. Lewis et al. report that the number of documents retrieved affects both performance and running time. Passing more reduces misses, but as the RAGAS paper notes, overly long context is harder for an LLM to use well and costs more to process.
Set the answering instructions
- Answer only from the material provided, and say that nothing was found when it does not cover the question
- Cite sources (document name, version, page or clause number) so users can check the original
- Fix the output structure (answer, basis, caveats) and check in code that required fields are present
Under LLM07 (Misinformation), the OWASP Top 10 for LLM Applications 2026 likewise recommends grounding claims in authoritative, current sources and requiring structured outputs with mandatory fields to prevent omissions.
RAG evaluation: measure retrieval and generation separately
Measure RAG accuracy in two parts: did the system retrieve the documents it needed, and did it answer the question faithfully from them? Judging only whether the answer was good leaves you unable to tell whether a failure came from retrieval or generation, and so unable to decide what to fix.
The RAGAS paper, which proposed a framework for automated evaluation of RAG pipelines, centres on three quality dimensions:
- Faithfulness: whether the answer is grounded in the context provided (the retrieved documents)
- Answer relevance: whether the answer addresses the question that was actually asked
- Context relevance: whether the retrieved context is focused on the question, with as little irrelevant information as possible
For retrieval itself, measure how often the document or chunk holding the answer appears among the top results. For business use, also check that citations are correct (the cited passage really says what the answer claims) and that the system says it found nothing when the documents do not cover the question.
Building the evaluation set
- Collect real questions: from enquiry logs and help-desk records, pick frequent questions and those where a wrong answer would be costly
- Record the ground truth: for each question, have the business owner note the source document, the passage and a model answer
- Include hard cases: questions the documents do not cover, questions whose answer changed between versions, and questions that span several documents
- Re-run on every change: whenever documents, chunking, the embedding model, the LLM or the prompt change, compare results on the same set
NIST AI 600-1 advises against extrapolating a generative AI system's performance from narrow, non-systematic and anecdotal assessments, and calls for reviewing and verifying the sources and citations in outputs both before deployment and in ongoing monitoring. Rather than trying a handful of questions and going live, record results and agree pass criteria in advance; that is what lets you explain the system's quality.
Automated evaluation with an LLM as the grader covers many questions quickly, but the grader can be wrong too. Combine automated scores with spot checks by the business owners.
RAG security: permissions, indirect prompt injection and the vector database
What sets RAG security apart is that the indexed documents become both a new input channel into the LLM and a new channel for leaks. The OWASP Top 10 for LLM Applications 2026, published in August 2026, sets out RAG-related risks and mitigations under LLM01 (Prompt Injection), LLM02 (Sensitive Information Disclosure) and LLM09 (Vector and Embedding Weaknesses), among others.
| Risk | What happens | Main mitigations |
|---|---|---|
| Access beyond permissions (LLM02, LLM09) | HR, contract or customer chunks the user may not see are retrieved and appear in the answer | Enforce document- and chunk-level authorisation inside the index query; separate indexes for highly sensitive areas |
| Indirect prompt injection (LLM01) | The LLM follows instructions hidden in an ingested document, and the answer is manipulated | Validate documents before ingestion, strip hidden and white-on-white text, index external documents separately |
| Knowledge base poisoning (LLM01, LLM09) | A crafted document is retrieved for a target question and steers the answer | Limit ingestion sources, record provenance and ingestion time, have people review external content |
| Embedding inversion (LLM09) | Original text is reconstructed from leaked vectors or backups | Classify and encrypt the vector database and its backups at the same level as the source documents |
| Leaks through logs (LLM02) | Observability tools record full prompts, answers and retrieved chunks | Restrict access to logs and scrub sensitive data before it is logged |
On permissions, OWASP is explicit. Similarity search does not respect access control lists, and a chunk already passed to the model cannot be taken back, so filtering in the application after retrieval is too late. Authorisation belongs before retrieval, inside the index query, and highly sensitive workloads call for physically separate indexes per tenant or trust zone. A mostly public document can contain a confidential paragraph, so the unit of control is the chunk.
IPA's guideline likewise lists, as a RAG problem, that confidential data stored in the vector database becomes visible to every user, and describes setting which users may see information according to its sensitivity. In interviews IPA held with organisations between March and May 2024, many took applications from departments or projects and stored each one's documents in a separate area.
On indirect prompt injection, OWASP cites research (Zou et al., 2025) in which as few as five poisoned documents reached roughly 90% attack success against a knowledge base of millions of texts. Even an internal system opens the same channel if it ingests text written by outsiders, such as supplier files, job applications or customer emails. The appendix to the AI Guidelines for Business also asks providers to consider regular management and audits of the data RAG searches, to detect unauthorised changes early.
If the LLM does more than answer, for example sending emails or updating records, OWASP advises reading the Top 10 for Agentic Applications from the same project alongside this list.
Cloud API, private connectivity or on-premises (local LLM)
There are three broad options, depending on where the LLM and the index run. The deciding factors are how sensitive the documents are and whether they may leave the organisation, usage volume, the model capability required, and whether you can run GPUs and models in-house.
| Aspect | Cloud API | Cloud over private connectivity | On-premises or local LLM |
|---|---|---|---|
| Set-up | Call an LLM API over the internet | Use a cloud LLM through a leased line or private connection | Run the model in your own data centre or on your own devices |
| Where data goes | Questions and retrieved text go to the provider | Data goes to the provider, but not over the internet | Data stays inside the organisation |
| Models available | Choose from the provider's large models | Choose from models in services that support private connectivity | Limited to models that fit your GPUs |
| Getting started | Fast, with low upfront cost | Network and environment set-up needed | GPU servers to buy and build |
| Operating burden | Low (the provider runs the model) | Moderate (connectivity and environment) | High (updates, patching and monitoring are yours) |
| Check before choosing | Whether inputs are used for training or retained, and where data is stored | Network path, where logs are kept, the provider's certifications | GPU capacity, update arrangements, model licences |
Even with a cloud API, minimise what you send. OWASP recommends sending external providers only the fields a task needs, and enforcing no-training and no-retention technically rather than relying on policy text alone.
A local LLM, run in your own environment, keeps documents and questions in-house, but your GPUs limit the size of model you can use. Among models strong in Japanese, LLM-jp, a project hosted by the large language model research and development centre at Japan's National Institute of Informatics (NII), aims to build open models strong in Japanese and publishes them. NTT began offering tsuzumi 2 in October 2025, a lightweight model developed to run on a single GPU with 40 GB of memory or less. Check the licence and terms of use for each model.
Whichever option you choose, decide where the index, its backups and the logs live by the same standard as the LLM. Because embeddings can be inverted to recover the original text, OWASP asks that vector database backups be treated at the same sensitivity level as the source documents. Putting only the LLM in a closed network does not keep data in-house if the index or the logs sit in an external service.
A private or on-premises deployment does not make a system secure by itself. Permission design, log management, patching and model updates are needed in every option.
Operating RAG: document updates, monitoring and re-evaluation
A RAG system is not finished at launch; it is a continuing cycle of updating documents and re-evaluating. IPA's guideline says that keeping RAG accurate requires updating the vector database regularly and keeping it as current and free of duplicates as possible.
- Automate updates: connect the index to where the authoritative documents live so revisions and withdrawals show up promptly, and delete the embeddings of deleted documents within a set time
- Re-embed everything when you change the embedding model: OWASP advises against mixing old and new vectors
- Keep immutable retrieval logs: record which question, under whose permissions, returned which chunks
- Collect user feedback: gather ratings on answers and the questions the system could not answer, and use them to add documents and update the evaluation set
- Re-measure regularly: run the evaluation set before and after any change to models or settings
Explaining the system to users is part of running it. IPA's guideline cites misconceptions such as 'RAG can answer any question correctly' and points to the gap between what users expect and what RAG can do. The appendix to the AI Guidelines for Business cautions that RAG tends to make answers converge, so it may not suit work that needs diversity or originality.
Steps to build a RAG system
Build the first RAG system small, prove it with evaluation, then widen it. A typical sequence:
- Define the use and scope: decide whose questions it answers and about what, which documents are in scope, and what it must not answer
- Inventory data and risks: map the authoritative copies, versions, classifications and access rights, and choose the deployment option by whether data may leave the organisation
- Build the evaluation set first: prepare real questions and their ground truth with the business owners, and agree pass criteria
- Prototype and measure: assemble ingestion, chunking, retrieval and prompts, and improve them against the evaluation set
- Test security: before launch, test for access beyond permissions, indirect prompt injection and the handling of logs
- Release to a limited group, then widen: open it to a few departments and fix documents and settings from user feedback and retrieval logs before extending it
The AI Guidelines for Business (Ver1.2) include a checklist and worksheet (Appendix 7) for reviewing your measures, which help confirm what is expected of you as a developer, provider or user of AI.
How findn can help
As part of its System Development service, findn builds LLM applications that answer from company data (RAG) and AI agents, covering model selection and customisation including Japanese LLMs, private and on-premises deployment, evaluation, guardrails and LLMOps. The work spans requirements, design, testing and operations, under an ISO/IEC 27001:2022 certified management system.
Questions and answers
- Should we use RAG or fine-tuning?
- Start with RAG if you want answers based on your organisation's knowledge: updating knowledge only means changing the index, and answers can cite their sources. Consider fine-tuning when you need to change behaviour rather than knowledge, such as a consistent style or output format. The two can also be combined.
- Does RAG eliminate hallucinations?
- No. RAG reduces hallucinations, but IPA's guideline states that using RAG cannot remove the risk of hallucination. Combine instructions to say when nothing was found, visible sources, continuous measurement on an evaluation set, and human review of important decisions.
- Can a RAG system run fully on-premises?
- Yes. If the LLM, the embedding model, the vector database and the logging platform all run in your own environment, documents and questions never leave the organisation. In exchange, your GPUs limit the size of model you can use, and model updates and vulnerability management are yours. Options include Japanese models such as NTT's tsuzumi 2, designed to run on a single GPU, and the openly published LLM-jp models from NII.
- How long does it take to build a first RAG system?
- There is no standard duration. It depends on the volume and condition of the documents (how many are scanned PDFs, whether versions are controlled), the complexity of the permission design, the deployment option (including GPU procurement for on-premises) and the accuracy required. Limiting the scope to one business process and building the evaluation set before the prototype makes estimates easier, and lets the decision to go live rest on numbers.
- How do you measure RAG accuracy?
- Measure retrieval and generation separately. For retrieval, check whether the document holding the answer is among the top results and whether the retrieved context is focused on the question (context relevance). For generation, check whether the answer is grounded in the retrieved documents (faithfulness) and whether it answers the question (answer relevance). Use an evaluation set built from real business questions and re-run it after every change.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks Opens an external site (arXiv (Lewis et al., NeurIPS 2020))
- Ragas: Automated Evaluation of Retrieval Augmented Generation Opens an external site (arXiv (Es et al.))
- OWASP Top 10 for LLM Applications 2026 Opens an external site (OWASP GenAI Security Project)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1) Opens an external site (National Institute of Standards and Technology (NIST), US)
- Guideline for Introducing and Operating Text-Generation AI (Japanese) Opens an external site (Information-technology Promotion Agency (IPA), Japan)
- AI Guidelines for Business, Ver1.2 (Japanese page with provisional English translation) Opens an external site (Ministry of Internal Affairs and Communications (MIC) and Ministry of Economy, Trade and Industry (METI), Japan)
- LLM-jp Opens an external site (National Institute of Informatics (NII), Japan)
- NTT Large Language Model tsuzumi 2 (Japanese) Opens an external site (NTT R&D)
