How agents use knowledge
Beyond Presence extracts the text of each file page by page, stores it in full, and builds a search index for it. During a conversation, the agent receives your knowledge in one of two ways:- In the prompt: If the text of all attached files fits in the agent’s prompt, Beyond Presence adds the full text to the agent’s instructions. This is the case for most small knowledge bases.
- Through retrieval: If the text no longer fits in the prompt, the agent gets a
search_knowledgetool instead. The agent searches the knowledge base before answering, using a hybrid of semantic (vector) and keyword (BM25) search. Results are labelled with the document name and page, so the agent can cite its source. If a search finds nothing relevant, the agent says it does not have that information instead of guessing.
File lifecycle
Every knowledge file has astatus:
Text files and submitted PDF files both come back as
processing.
Indexing a long document can take several minutes.
Poll Retrieve Knowledge File until the status is available before you attach the file to an agent.
Attaching a file that is still processing or has failed returns an error.
Create and wait for a text file
Python
format: "pdf" and upload_num_chunks, upload each chunk with Upload Knowledge File Chunk, and finish with Submit Knowledge File. Then poll the file the same way.
Extracted text
Anavailable file includes the extracted text and an is_text_complete flag.
For large files, text contains only the start of the file and is_text_complete is false.
The agent still searches the full file either way. The flag only describes how much text the API returns.
Failed files
When a file has thefailed status, read its reason to find out what to change.
Common reasons include:
- Scanned PDF: No text could be extracted from any page. The PDF is most likely a scan or image-only file. Run it through OCR and upload it again.
- Too many pages: The PDF has more than 1,500 pages. Split it into smaller files and upload them separately.
- Too large to index: The file contains more text than one index can hold. Split it into smaller files and upload them separately.
- Unreadable or damaged PDF: The file is corrupt, is not a PDF, or stopped parsing partway through. Re-export it and upload it again.
- Unexpected error: Something went wrong during indexing. Upload the file again to retry.
Limits
The same 250,000-character limit applies in the Studio and the API.