# How Large Language Models Retrieve and Use Information

> Canonical source: [How Large Language Models Retrieve and Use Information](https://www.answer.cloud/pages/how-large-language-models-llms-retrieve-data-hosted/)

Distinguish model knowledge, search tools, retrieval systems, uploaded documents, and connected data.

An LLM does not use one universal information pipeline. What an answer draws on depends on the product, its tools, the available context, and the question.

## Four common sources

1. **Existing model knowledge:** learned patterns and information from training. Publishing a new website page does not immediately update those model parameters.
2. **Web search and page retrieval:** a search-enabled assistant can find public sources and fetch relevant pages.
3. **Uploaded or connected documents:** a product may use files, organization data, or other authorized sources.
4. **Conversation context:** information supplied by the user or already included in the current interaction.

## One possible retrieval design

A retrieval-augmented system can index documents, search for relevant passages, and supply selected material to the model. It may use keywords, embeddings, reranking, or several searches. Passage size and selection vary; some products also process entire documents or larger contexts. A fixed 100–300-word chunk is not a universal rule.

## What website owners can improve

Publish clear, accurate HTML with descriptive titles and headings. Keep URLs stable and important pages accessible. Use reliable sources and update stale claims. A clean text version can help tools that use it, while the normal web page remains important.

Neither crawl access nor a text mirror guarantees that a particular answer uses your page. Measure retrieval requests and sampled citations separately with the [visibility metrics guide](https://www.answer.cloud/pages/llm-visibility-metrics-the-kpi-playbook-for-seo-and-marketing-leaders-hosted/).