Knowledge Index

Last modified by Thomas Mortagne on 2026/10/09 15:24

This document describes the implementation of the knowledge index, see the Planned Features and Architecture Overview for an overview of the knowledge index.

Data Model

As explained in the overview, the knowledge index is structured into collections of documents. We identified indexing arbitrary data like PDFs as an interesting use case so we plan to support uploading arbitrary file types. As in XWiki itself, Apache Tika is used to extract their content.

To simplify right management of collections, a simplified view of rights is provided based on a list of allowed users and groups for each right with no way to deny rights. Leaving a list empty means nobody has the right. Saving a collection that would deny the right to manage the collection to the current user is prevented. We identified the following rights for collections:

  • Right to use/view a collection for RAG (individual documents could still be subject to further restrictions)
  • Right to manage a collection's content (for adding and updating content, e.g., via the REST API)
  • Right to manage a collection's settings

The following shows the required fields we identified for collections and documents:

WAISE Collections:

  • Name
  • Chunking method
  • Embedding model
  • Users and groups who can use the collection
  • Users and groups who can manage the content of the collection
  • Users and groups who can administrate the collection
  • Right check method for retrieval
  • Parameter string for the right check method (usually the URL)
  • XWiki space for storing the documents help
    • Maybe for indexing documents already stored in WAISE. But then we don't have the same data on them, multiple attachments, ... - so maybe just duplicate content in this case?

WAISE Documents:

  • Id
  • Title
  • Language
  • URL
  • Content (can be binary)
  • Mime type

Storage

Before indexing the content, it is stored as regular XWiki documents. This makes it easy to leverage existing XWiki access control and App Within Minutes for the management of the content. Both WAISE collections and WAISE documents are identified by XObjects that also store some of the data.

By default, documents are stored in the following location: AILLMApp.Collections.{id}.Documents.{document_id}. Maybe, this can be changed by setting the space property of the collection to a different space.

Rights are mapped as follows:

  • Right to use/view a collection - might not be reflected in actual XWiki rights or view right on space
  • Right to manage a collection's content - edit right on the documents space, allows to see all documents
  • Right to manage a collection's settings - edit right on the space of the collection

A listener updates the configured rights upon saving a collection. If the collection has a custom space configured, content view and edit right aren't updated automatically.

Note that if view right is granted on the space of the collection's document, we need to actually check rights for regular view access in XWiki based on the right check method for retrieval. TODO: Figure out how complicated that would be. Note that would also mean that we would definitely cache access rights in XWiki which might or might not be desirable.

To store a document, the following mapping is applied from WAISE document properties to XWiki document:

  • Title -> Title
  • Language -> Language
  • Id -> XObject
  • URL -> XObject
  • Mime type -> XObject
  • Content -> content if text, attachment otherwise + Tika-extracted content in content

Indexing

Whenever a document is updated or added, it is added to an indexing queue.

Indexing worker thread takes documents from queue, chunks them, computes embedding vectors and adds them to the vector store (Solr).

For every chunk, store at least the document id and the language (for querying by language) and the position of first and last byte or character in the document, maybe also a chunk index (numbered from 0). We could also store the document reference where the collection's document is actually stored to retrieve metadata from there.

For the implementation, take inspiration from Solr indexer in XWiki. Also consider the case of re-indexing a whole collection!

Wild ideas:

  • Ask an LLM to generate an intro that puts the chunk into context, giving the LLM the original document + the chunk. If the original document is too big, summarize incrementally.

Query for a Specific Collection

  1. Embed query
  2. Find similar chunks
  3. Ask authorization method of the collection if documents are allowed
  4. Possibly augment chunks with extra context (chunks before/after)
  5. (Optional) re-ranking
  6. (Optional) Reorder chunks for optimal RAG results (e.g., combine chunks of the same document and put them in the correct order) and put most relevant chunks in the positions where the LLM pays attention. See, e.g., On Position Bias in Summarization with Large Language Models.

Here’s how DenseVectorField should be configured in the schema:

   - <fieldType name="knn_vector" class="solr.DenseVectorField" vectorDimension="1024" similarityFunction="cosine"/>
   - <field name="vector" type="knn_vector" indexed="true" stored="true"/>

REST API

The REST API uses the regular XWiki authentication or the custom Authentication for WAISE, so users are always executed and standard XWiki right checks can be used.

  • collections: /collections/:id/ (for get/delete/modify/dispose cache)
    • list (index) with/without metadata? with pagination!
    • get collection by id (document name)
    • create collection
    • delete collection - requires admin on collection
    • modify collection
    • dispose cache of access rights?
    • index status (like queue size)
    • documents: /collections/:id/documents/:id/ Requires manage collection content right!
      • list with pagination!
      • create
      • delete
      • update
      • get
      • dispose cache of access rights?
      • index status? (is document queued for indexing)

REST API synchronously stores the documents as XWiki documents, indexing is triggered asynchronously as described in the sections above.

Implementation Plan

  1. Create UI and API Maven module for indexing
  2. Create XClasses using AWM
  3. Create Java API for collections and documents - possibly with and without authorization, or add authorization check methods?
  4. Create REST API
  5. Create indexing queue + worker + actual index Java API

Get Connected