I’m a beginner who discovered Hugging Face a few days ago and I’m really impressed by what we can do here.
I was wondering if it’s possible to replicate the “domain search” feature (like in HuggingChat) for my own custom chatbots, essentially using it as a RAG approach.
Is there a straightforward way to crawl or connect data from a website URL for that purpose? If so, could you please explain how or point me to any relevant tools or examples?
Hello. There are two broad methods. One is to process the results of a normal web search using a programming language such as Python and pass the results to LLM yourself. The other is a method called Function Calling, in which you instruct LLM to execute a search tool and return the results. (There are various names for this method.)
In the case of the former, there are various useful libraries, so you should try searching for them. If you can use the latter, it is usually built into the LLM execution environment, so it is often found somewhere in the documentation.
Hi! One approach that works well is to build a structured knowledge base from your website instead of relying only on raw page crawling. If your site already has FAQs, service pages, pricing information, and educational content, you can chunk that content, generate embeddings, and store it in a vector database for RAG.
For local business websites, users often ask the same question in different ways, such as:
How much does junk removal cost?
Can I recycle old electronics?
Where should I dispose of a mattress?
What items are accepted for pickup?
A combination of semantic search and keyword search usually returns more reliable results than using only one approach, especially for conversational queries.
I’m also interested to know what tools others are using to crawl websites and keep the knowledge base updated automatically.