Skip to main content
A knowledge base is the set of knowledge files attached to an agent. The agent uses them to answer questions about your products, policies, or other domain-specific content. You can manage knowledge files in the Studio or through the Knowledge Files API.

How agents use knowledge

Beyond Presence extracts the text of each file page by page, stores it in full, and builds a search index for it. During a conversation, the agent receives your knowledge in one of two ways:
  • In the prompt: If the text of all attached files fits in the agent’s prompt, Beyond Presence adds the full text to the agent’s instructions. This is the case for most small knowledge bases.
  • Through retrieval: If the text no longer fits in the prompt, the agent gets a search_knowledge tool instead. The agent searches the knowledge base before answering, using a hybrid of semantic (vector) and keyword (BM25) search. Results are labelled with the document name and page, so the agent can cite its source. If a search finds nothing relevant, the agent says it does not have that information instead of guessing.
Beyond Presence picks the mode for you at the start of each conversation. You don’t need to configure it.
Realtime LLMs can’t use knowledge files. The API rejects creating or updating an agent that combines a realtime model with knowledge_file_ids. Use a non-realtime model, or detach the knowledge files from the agent.

File lifecycle

Every knowledge file has a status: Text files and submitted PDF files both come back as processing. Indexing a long document can take several minutes. Poll Retrieve Knowledge File until the status is available before you attach the file to an agent. Attaching a file that is still processing or has failed returns an error.

Create and wait for a text file

Python
For PDF files, create the file with format: "pdf" and upload_num_chunks, upload each chunk with Upload Knowledge File Chunk, and finish with Submit Knowledge File. Then poll the file the same way.

Extracted text

An available file includes the extracted text and an is_text_complete flag. For large files, text contains only the start of the file and is_text_complete is false. The agent still searches the full file either way. The flag only describes how much text the API returns.

Failed files

When a file has the failed status, read its reason to find out what to change. Common reasons include:
  • Scanned PDF: No text could be extracted from any page. The PDF is most likely a scan or image-only file. Run it through OCR and upload it again.
  • Too many pages: The PDF has more than 1,500 pages. Split it into smaller files and upload them separately.
  • Too large to index: The file contains more text than one index can hold. Split it into smaller files and upload them separately.
  • Unreadable or damaged PDF: The file is corrupt, is not a PDF, or stopped parsing partway through. Re-export it and upload it again.
  • Unexpected error: Something went wrong during indexing. Upload the file again to retry.

Limits

The same 250,000-character limit applies in the Studio and the API.