When to update
- After a meaningful change to your site: new pricing, new hours, a new product line, a changed policy.
- When Test a question returns an old passage.
- When conversations show the agent quoting something that’s no longer true.
Re-read a website
There’s no separate re-crawl button. You re-read a website in one of two ways:- Save a change to its exclusions. Changing which pages are left out (below) and clicking Save & re-read the site reads the whole site again. The agent keeps answering from the current pages until the new read is finished.
- Remove the source and add the address again. This is a full fresh crawl. The agent has no knowledge from that site until the crawl finishes, which is usually a minute or two.
Exclude pages from a website
Plenty of what a crawl reaches is noise the agent should never answer from: sign-in pages, terms and conditions, blog archives, event listings, expired promotions. Leaving them out also gives the pages you do care about more room within the crawl’s limits.1
Open the website source
On Knowledge, click a ready website source. You’ll see its pages grouped by the first part of the path, such as
/blog or /services, with each page’s size.2
Find the pages
Use Search pages to filter by path or title. The sort button switches between sorting by path and sorting by size, which makes one huge page easy to spot.
3
Untick what should be left out
Untick a single page, or untick a group to leave out everything under that path. Excluded paths are listed under Left out of the knowledge base.
4
Save
Click Save & re-read the site. It takes a minute. The agent answers from the current pages until the new read finishes. Once it has, the excluded pages are gone from the page list and from the agent’s knowledge.
- A path covers everything beneath it. Excluding
/toolsalso excludes/tools/pricing-calculator, but not/toolsmith. - Case, trailing slashes, query strings and
#fragments are ignored. - Exclusions stick. They’re saved on the source and applied every time the site is read.
- Excluded pages are never fetched. The page and text limits they would have used go to the rest of the site.
- You can exclude the page you entered. The crawler still reads it to find links to other pages, but doesn’t index it.
- You can’t exclude the whole site (
/). To stop using a site, remove the source. - Up to 200 excluded paths per source.
Remove a source
Click the bin icon on the source’s row. The dialog asks Remove ”…”?. Choose Remove to confirm or Keep it to cancel. The source’s passages come out of the knowledge base straight away, and the agent stops answering from them. There’s no undo: to get it back, you add the website, file or text again.Stuck or interrupted indexing
Websites and files are indexed in the background. If a source still shows “crawling the site…” or “processing…” after 30 minutes, WireDesk treats the job as interrupted:- A website may be tried again automatically, up to three attempts. After that it fails with “this site could not be indexed after 3 attempts — it may be too large or too slow to crawl”. Excluding large sections you don’t need usually helps.
- A file or a note can’t be retried, because WireDesk doesn’t keep the original. It fails with “indexing was interrupted — add this source again”.
- You may also see “indexing was interrupted — remove this source and add it again”.
A regular routine
- Weekly: read a sample of conversations. Every “I don’t have that to hand” is a question your knowledge doesn’t answer. Write a note for each one.
- After a site change: re-read the site, then check two or three affected questions in Test a question.
- Occasionally: open each website source, sort by size, and exclude anything the agent doesn’t need.