Knowledge sources
Eighteen questions from a university network
A university network tested our chatbot for a week and sent eighteen follow-up questions. Fourteen were bugs on our side. What we built from them.
Contents
Two weeks ago a university network started feeding our chatbot its complete website in a test project: around a thousand pages, more than a hundred programmes across many universities, a course finder with filters and page numbers. The bot is not public there; it runs only in the test chat. The person in charge still tested it the way prospective students would ask, not with a test protocol. The result: eighteen follow-up questions, spread over three emails.
Fourteen of them were bugs or missing features on our side. That is the honest part of this post. The other part is what we built from them in one week. All of it is in the product for every customer today.
The crawler only saw the sitemap
The first question was the hardest: “The chatbot doesn’t use the information from the course finder.” Correct. Our crawler had followed the sitemap, and the sitemap listed 174 pages and 269 press releases, but not a single programme page. The university’s content management system lists programmes through an extension that never appears in any sitemap. Google doesn’t notice, Google follows links. We didn’t.
Since then the crawler follows links in addition to the sitemap. Two weeks later came the sequel: one programme was still missing. The course finder spreads its result list over seven pages, and pages two to seven are only reachable through addresses with parameters. Exactly the kind of address we deliberately skipped, because on many websites they stand for filter combinations that show the same list in a hundred variations. For this network, thirteen programmes hung on them alone. Now the crawler tells plain page numbers from filters: it follows page numbers, collects the links and does not store the list page as content.
A passage that doesn’t know which page it belongs to
“I’m still not happy with the answers about fees.” The fees were on the page and in the knowledge base, verbatim. Still the bot said it could find nothing. The reason was as simple as it was invisible: for search, pages are split into passages of about 600 characters. The fees passage said “the module fee is 65 euros” but not for which programme. That was only in the page title. And more than a hundred programme pages had an almost identical fees passage. So the search found the introduction of the right programme and the fees passage of a wrong one.
Since then every passage carries its page title as the first line. It sounds trivial. It is the single change with the biggest effect from this round, and it now applies to uploaded texts and PDFs as well.
”Doesn’t exist” is not an answer
When that one programme was missing from the index, the tester asked the bot directly. It replied: “No, this programme is not offered at this university.” That was wrong, and worse than an “I don’t know”: it sounded like a verified statement. Not finding something is no evidence that it does not exist. The bot may no longer claim that, unless a source says so explicitly.
The same question produced the page list as it is today: every crawled page with status, number of passages and error reason. Above it, a field where you enter a missing address. The page is fetched immediately and stays pinned, even if no link leads there any more.
Who is actually responsible?
Two questions were about people. First the bot quoted the direct line of a staff member it had found on some page, for questions she was not responsible for. So we built central contact details: once the customer enters them, the bot names only those and quotes no phone number from the pages any more.
Then the opposite direction: every programme page has a block “Your contact persons” at the bottom, and exactly those should be named, not the switchboard. That became a switch. When it is on, the bot may name people with their direct line if a page explicitly lists them as contact persons with a role. Someone merely mentioned in the text, such as the academic lead, does not become a contact.
And the third question of this kind: “How do I teach the chatbot that we are not a university?” The network is a state institution that coordinates distance-learning programmes. Asked for directions to a campus, the bot recommended calling “the university” and gave the network’s own number. It simply did not know whom it represented. Its prompt only stated its own name. Today every bot has a field “About the organisation”, and a rule never to guess what kind of organisation stands behind it.
The smaller things
- A blocked word in the bot configuration silently blocked 551 passages. The interface now shows, for every blocked term, how many passages it would hit.
- The bot switched between formal and informal address within one conversation. The form of address is now part of the instruction and follows the tone setting.
- The answer language followed the widget’s language setting, not the language of the question. Now it follows the visitor’s last message.
- Abbreviations like “MAPS” were never found because the model saw maps in them. There is now a glossary that applies before the search.
- Links in answers were a recommendation to the model, not a requirement, and nobody checked whether a link came from the sources at all. Now it is a requirement, and every link is checked against the sources actually used.
- Twelve programme pages took five to seven seconds to be served. Our crawler waited six. Now it waits twelve.
What we take from it
No test protocol would have found these eighteen points. They came from someone who knows their own content and asks questions you cannot come up with from the outside. The second lesson: almost every point became a rule or a switch, not a special case for one customer. And the third: a chatbot is only as good as your view of what it knows. That is why most of this week’s work went not into the model, but into the page list, the open questions and the sources under every answer.
If you have hundreds of pages behind a finder yourself: Chatbot for universities & continuing education.
Your own chatbot, built from your content
Add your website, upload PDFs or connect a folder. The bot answers only from your sources. Try it free for 7 days.