How to Train an AI Chatbot on Your Own Website Content
Pointing a crawler at your homepage is the easy part. This is what actually determines whether the answers come back accurate, and how to fix them when they do not.

The short answer
- Training a chatbot on your website means crawling your pages, splitting them into passages, converting those into embeddings, and retrieving the closest ones at question time. Answer quality is decided at the splitting and retrieval stages, not by the model.
- Crawl breadth is the most common mistake in both directions: too few pages and the bot cannot answer, too many and pricing pages compete with blog archives for the same question.
- Pages that render entirely in JavaScript, or PDFs that are scans, contain no text a crawler can read — they index as empty and produce confident silence.
- Budget for a review pass after go-live. The first version of any knowledge base is wrong in ways only real questions reveal.
What training on your website actually means
It is worth being precise, because the word training suggests something that is not happening. Your content is not used to retrain a language model. Nothing you upload changes the model's weights.
What happens instead is retrieval. Your pages are fetched and split into passages of a few hundred characters. Each passage is converted into an embedding, a numeric representation of its meaning, and stored. When a visitor asks a question, that question is embedded the same way, the closest passages are retrieved, and those passages are handed to the model along with the question and an instruction to answer only from what it was given.
This distinction matters practically. It means answer quality is governed by what got retrieved, not by how clever the model is. A better model cannot rescue a passage that was never found. It also means you can fix a wrong answer in minutes by fixing the source content, rather than waiting on anything to retrain.
Choose what to crawl, and what to leave out
The instinct is to point the crawler at the homepage and let it find everything. This is where most setups go wrong, in both directions.
Crawl too narrowly and the bot cannot answer ordinary questions, because the page holding the answer was never indexed. Crawl too broadly and you create competition: a three-year-old blog post that mentions pricing in passing now competes with your actual pricing page for the question about pricing. Retrieval picks the closest passage, not the most authoritative one.
- Include the pages a customer would be sent to by a good salesperson: product and service pages, pricing, FAQ, policies, shipping and returns, contact details, opening hours.
- Include documentation and setup guides if you have them. These answer the highest-volume support questions.
- Exclude blog archives, press releases and news posts unless they are evergreen. They age badly and are the single most common source of confidently outdated answers.
- Exclude careers pages, legal boilerplate you do not want quoted conversationally, and anything behind a login.
- Exclude old landing pages and superseded campaigns. If a page contradicts your current pricing, remove it from the crawl before you debug anything else.
The content problems that produce empty answers
A crawler reads text. If your page does not contain text at fetch time, it indexes as an empty document — and an empty document produces a bot that says it does not have that information, while your dashboard cheerfully reports that the page was synced successfully.
Three causes account for almost all of these cases.
- JavaScript-rendered pages. If the content is injected by a framework after load, a simple fetch sees an empty shell. Test by viewing the page source rather than the rendered page — if you cannot find your paragraph text in it, neither can the crawler.
- Scanned PDFs. A PDF made from photographs of pages has no text layer. It needs OCR before it holds anything indexable.
- Text baked into images. Pricing tables and comparison charts saved as PNGs are invisible to retrieval, however clearly a human reads them.
When a page cannot be crawled, pasting the text directly into the knowledge base is the reliable fallback — it skips the crawler entirely and indexes the same way.
See how knowledge sources workWrite content that retrieves well
Content written for a chatbot is not different from content written well for people, but a few habits matter more than usual because passages are retrieved out of context.
The critical constraint is that a passage may be read alone. A paragraph beginning "As mentioned above, this applies to all three tiers" is useless once separated from what was above it. Self-contained paragraphs retrieve far better than flowing ones.
- Put the answer in the first sentence of the section, then explain. A passage that opens with the answer is a passage that can be quoted.
- Use the words your customers use, not your internal vocabulary. If they say delivery and you say fulfilment, index both.
- Keep one topic per section under a descriptive heading. Headings survive the splitting process and help the right passage surface.
- Repeat essential qualifiers rather than referring back. Restating that a policy applies to UK orders costs one clause and prevents a wrong answer.
- State facts explicitly, including things that feel obvious. Opening hours and refund windows are asked constantly and are often nowhere on the site in plain text.
Test before you go live
The temptation after a successful crawl is to publish immediately. Twenty minutes of testing prevents most complaints.
- 1Write down the twenty questions your team answers most often. Pull them from your inbox rather than from imagination.
- 2Ask each one exactly as a customer would type it, including typos and shorthand.
- 3Ask five questions you know the site does not answer. The bot should decline rather than improvise — this is the single most important test.
- 4Ask three questions where the honest answer is that it depends. Check whether the bot asks a clarifying question or guesses.
- 5Ask about something you deliberately removed from the crawl, to confirm the exclusion worked.
- 6Fix the source content for each failure, re-index, and repeat until the failures are only genuine escalations.
Keep it accurate after launch
A knowledge base is not a project with an end date. It drifts because your business changes and the content does not.
The highest-value habit is reviewing the questions the bot could not answer. That list is the most honest description of the gap between what customers want to know and what your website says. It is also, usually, a list of pages you should have written anyway.
Re-index whenever you change pricing, policies or opening hours, and set a recurring reminder to re-crawl on a schedule so a quietly edited page does not sit stale for months.
Frequently asked questions
How long does it take to train a chatbot on a website?
Crawling and indexing a typical small business site of 20 to 50 pages usually takes a few minutes. The work that determines quality is the review pass afterwards — testing real questions and fixing the content behind wrong answers — which realistically takes two to four hours for a first deployment.
How many pages should I crawl?
Quality matters far more than quantity. Twenty accurate, current, well-structured pages will outperform two hundred pages that include outdated blog posts and superseded landing pages. Start with the pages a good salesperson would send a customer to, then add sources as real questions expose gaps.
Why does my chatbot say it has no information when pages synced successfully?
Almost always because the pages indexed as empty. The most common causes are JavaScript-rendered content that is not present in the raw HTML, scanned PDFs with no text layer, and information that only exists inside images. Check the page source rather than the rendered page — if your text is not there, the crawler never saw it.
Do I need to retrain the chatbot when my website changes?
You need to re-index, which is different and much faster. The underlying model is never retrained on your content. Re-index whenever pricing, policies, hours or product details change, and schedule a periodic re-crawl so edits made without telling anyone do not sit stale.
Can a chatbot be trained on PDFs and documents?
Yes, provided the file contains real text. A PDF exported from a word processor indexes well. A PDF that is a scan or photograph of a page contains no text layer and needs OCR first, otherwise it indexes as an empty document.


