Vector Stores
The vector store is the persistence layer for embedded chunks. Rag.NET ships six implementations, each registered via a fluent extension method on RagBuilder. The interface is designed to be swapped without changing any pipeline code.
Feature matrix
| Feature | PgVectorStore | QdrantVectorStore | AzureAISearchVectorStore | WeaviateVectorStore | ChromaVectorStore | PineconeVectorStore | RedisVectorStore |
|---|---|---|---|---|---|---|---|
| Package | Rag.NET.VectorStores.PgVector | Rag.NET.VectorStores.Qdrant | Rag.NET.VectorStores.AzureAISearch | Rag.NET.VectorStores.Weaviate | Rag.NET.VectorStores.Chroma | Rag.NET.VectorStores.Pinecone | Rag.NET.VectorStores.Redis |
| Dense (semantic) search | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Hybrid search (native) | No — BM25 fallback | No — BM25 fallback | Yes (IHybridSearchable, fused score) | Yes (IHybridSearchable, fused score) | No — BM25 fallback | No — BM25 fallback | No — BM25 fallback (why) |
Sparse search (SPLADE, ISparseSearchable) | Yes (enableSparseVectors: true) | Yes (enableSparseVectors: true) | No | No | No | Yes (EnableSparseVectors = true) | No |
| Metadata filtering | Yes (JSONB @>) | Yes (payload match / numeric range) | Yes (typed metadata_entries/any(...)) | Yes (typed where on meta_* props) | Yes (where $eq/$and) | Yes (filter $eq/$and) | Yes, on declared keys only |
| Typed metadata round-trip | Yes (native JSONB types) | Yes (native payload types) | Yes (typed complex-collection slots) | Yes (typed auto-schema props) | Yes (native values; dates as sentinel) | Yes (native values; dates as sentinel) | Yes (typed JSON blob in a metadata hash field) |
ICollectionManageable | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Similarity function | Cosine (via <=>); dot product when sparse (<#>) | Cosine | Cosine | Cosine | Cosine | Cosine (dotproduct when sparse) | Cosine (distance converted to similarity) |
| Index algorithm | HNSW at ≤ 2000 dims, exact scan above (see below) | HNSW | HNSW | HNSW | HNSW | Serverless (managed) | HNSW |
| Persistence | PostgreSQL | Qdrant server | Azure managed | Weaviate server | Chroma server | Pinecone managed | Redis Stack / Redis 8+ |
Interface hierarchy
The sparse subtypes are opt-in registrations, not separate packages — the dense base type
deliberately does not implement ISparseSearchable, so a store is ISparseSearchable
capability probe is honest.
Shared interface
All six implement IVectorStore:
public interface IVectorStore
{
Task StoreAsync(IReadOnlyList<EmbeddedChunk> chunks, CancellationToken cancellationToken = default);
Task<IReadOnlyList<SearchResult>> SearchAsync(
ReadOnlyMemory<float> queryEmbedding,
SearchOptions options,
CancellationToken cancellationToken = default);
Task DeleteByDocumentIdAsync(string documentId, CancellationToken cancellationToken = default);
}
SearchOptions carries TopK, MinScore, and MetadataFilter. Hybrid routing is not part of it: the pipeline decides between SearchAsync and IHybridSearchable.HybridSearchAsync before calling the store (see Retrieval — How the hybrid path is selected), so a store implementation never sees a hybrid flag.
Native hybrid scores are ordinal, not similarities
Azure AI Search and Weaviate answer IHybridSearchable.HybridSearchAsync by fusing a keyword (BM25) ranking with a vector ranking inside the backend, and hand back the fused rank as Score. A fused rank carries no similarity meaning — only order — so both stores declare this through IHybridSearchable.HybridScoreScale (the interface's own default, ScoreScale.OpaqueRanking; neither store overrides it — see Score scale for the sibling declaration on the dense path), and neither applies SearchOptions.MinScore to that score: thresholding a rank as if it were a cosine similarity would keep or drop results arbitrarily.
This does not mean the retrieval pipeline filters a request's MinScore away. It means a request carrying one never reaches the native path to begin with: EnsembleBehavior.CanDispatchNatively requires MinScore to be exactly 0.0 (alongside no EnsembleOptions and no sparse arm), so a non-zero threshold is served by client-side RRF instead, where MinScore is honoured against the dense arm's real similarity score — see Retrieval — How the hybrid path is selected. The only caller who can still see an un-thresholded fused score is one invoking HybridSearchAsync directly, bypassing the pipeline — which is why the declaration lives on IHybridSearchable itself rather than being left for every such caller to rediscover.
Typed metadata
Chunk metadata values are typed (MetadataValue: string, number, boolean, or date — see the ingestion guide), and every store persists the type: a page written as the number 3 is stored as a number, read back as a number, and filtered as a number — on Redis, "filtered" holds only for declared filterable keys; every key is still stored and returned regardless. MetadataFilter takes the same typed values, so a numeric filter is a numeric comparison in the store, not a string match:
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance", // string match
["page"] = 3, // numeric match — NOT the string "3"
},
});
Filter values are kind-sensitive: filtering on the string "3" does not match a stored number 3. Metadata stored before values carried types reads back losslessly as string-kind values; see each store's section (and the Azure AI Search migration note) for what that means for filtering old documents.
Initialisation
You do not have to call InitializeAsync. Every store in this library creates its index or
collection on first use — the first StoreAsync or SearchAsync — so a first ingest against a
backend that has never been provisioned works without any startup step.
DeleteByDocumentIdAsync deliberately does not initialise: a delete against a collection that
does not exist has nothing to delete, so provisioning one to satisfy it would be waste — and on
pgvector an inline HNSW build under a write-blocking lock, set off by a delete.
It runs once per store instance, behind a lock, and only a successful run is remembered: a transient failure at startup does not leave the store permanently broken, the next call tries again.
InitializeAsync is on IVectorStore, so nothing needs casting to a concrete store type:
// Optional. Pays the cost at a moment you choose, rather than inside the first request.
var store = provider.GetRequiredService<IVectorStore>();
await store.InitializeAsync();
Call it explicitly when you want to:
- provision at startup rather than inside the first user-facing request. This matters most on
pgvector against a large existing table: initialisation is also the migration path, and it
builds the HNSW index inline while holding a
ShareLockthat blocks concurrent writes. Leave it implicit and that cost lands inside whichever request happens to be first; - fail fast on a misconfiguration at a point you control;
- re-create a collection you have just dropped with
DeleteCollectionAsync.
The trade-off, stated plainly. Creating on first use means a mistyped index or collection name silently creates a new empty one instead of failing. You get an empty search result rather than an error naming the missing index. If you would rather have the error, provision your backends out of band and treat any missing collection as a fault — the stores will still find the existing one and not create anything.
Dropping the bound collection through DeleteCollectionAsync resets this: the next operation
creates it again, rather than writing to something that no longer exists.
On pgvector, first use also migrates an older table — adding the unique
(document_id, chunk_index) key that upserts need. Where that is impossible, because the table
already holds duplicate rows under one key, the first StoreAsync fails with the same explanation
InitializeAsync would have given, naming the duplicate count and the query to inspect them. It
deletes nothing.
Implementing
IVectorStoreyourself?InitializeAsynchas a default no-op body, so an existing implementation keeps compiling and behaving exactly as before. Override it if your store provisions a backend.
Collection management
All six also implement ICollectionManageable, registered alongside IVectorStore in the DI container:
public interface ICollectionManageable
{
Task CreateCollectionAsync(string name, int vectorDimensions, CancellationToken cancellationToken = default);
Task DeleteCollectionAsync(string name, CancellationToken cancellationToken = default);
Task<bool> CollectionExistsAsync(string name, CancellationToken cancellationToken = default);
}
Resolve it directly from DI when you need to manage the index lifecycle:
var manageable = provider.GetRequiredService<ICollectionManageable>();
if (!await manageable.CollectionExistsAsync("rag-index"))
await manageable.CreateCollectionAsync("rag-index", vectorDimensions: 1536);
DeleteCollectionAsync is uniform across all six stores: deleting a collection that does not
exist is a no-op, so teardown code needs no exists-probe. The backends disagree underneath —
Weaviate and PgVector are idempotent themselves, while Chroma answers 404, Pinecone raises
NotFoundError, Qdrant reports result: false, and Azure AI Search answers 404 — and each
store absorbs its own flavour.
PostgreSQL + pgvector
Package: Rag.NET.VectorStores.PgVector
Stores chunks in a rag_chunks table. Uses the pgvector extension for dense search via the <=> cosine distance operator and, when sparse vectors are enabled, for SPLADE search via the <#> inner-product operator over a sparsevec column. Metadata is stored as JSONB with native JSON types — a number-kind MetadataValue is a JSON number, a boolean a JSON boolean, a date a {"$date": ...} wrapper — and filtered using PostgreSQL's containment operator, so a numeric filter is numeric containment.
Setup
services.AddRagNet(rag => rag
.UsePgVector(
connectionString: "Host=localhost;Database=ragdb;Username=postgres;Password=secret",
vectorDimensions: 1536));
The vectorDimensions must match the output dimension of your embedding model (text-embedding-3-small → 1536, mxbai-embed-large → 1024, etc.).
Add enableSparseVectors: true to register the sparse-capable subtype instead — see Sparse vectors (SPLADE) below.
Schema
InitializeAsync creates the following objects if they do not already exist:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE IF NOT EXISTS rag_chunks (
id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
document_id TEXT NOT NULL,
chunk_index INTEGER NOT NULL,
text TEXT NOT NULL,
metadata JSONB NOT NULL DEFAULT '{}',
embedding vector(<dimensions>) NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_rag_chunks_document_id ON rag_chunks (document_id);
-- The chunk key. StoreAsync upserts on it.
CREATE UNIQUE INDEX IF NOT EXISTS idx_rag_chunks_doc_chunk ON rag_chunks (document_id, chunk_index);
-- Dense ANN index — created only when <dimensions> <= 2000. See below.
CREATE INDEX IF NOT EXISTS idx_rag_chunks_embedding ON rag_chunks USING hnsw (embedding vector_cosine_ops);
These are created on first use, so no startup call is required — see Initialisation. To pay the cost at startup instead, which the HNSW note below is a good reason to do:
var store = provider.GetRequiredService<IVectorStore>();
await store.InitializeAsync();
InitializeAsyncis not always a quick startup step. Building the HNSW index over a large existing table is slow and memory-hungry, and it happens inline.CREATE INDEXtakes aShareLockonrag_chunks: concurrent writes block until it finishes, reads are unaffected. (CREATE INDEX CONCURRENTLYis deliberately not used — it cannot run inside a transaction, and a failed run leaves anINVALIDindex thatIF NOT EXISTSwould then skip forever.) Budget for it on the first run after upgrading.
vectorDimensionsis baked into the table at first initialize.CREATE TABLE IF NOT EXISTSmatches on the table name only, so pointing a store configured for 1536 at arag_chunkswhoseembeddingisvector(768)would skip the statement, report success, and keep the old typmod — after which every write and search fails on the mismatch, far from the setting that caused it.InitializeAsynctherefore fails fast, naming both dimensions: either construct the store withvectorDimensions:matching the table, or drop the table and re-ingest with the embedding model you are actually using. (Dropping discards every chunk stored; re-ingestion is what recomputes them.) The same check guardsCreateCollectionAsyncagainst itsvectorDimensionsargument.
Chunk key and upsert semantics
A chunk is keyed by (document_id, chunk_index), enforced by a unique index. StoreAsync upserts on that key, so re-storing a chunk replaces it rather than appending a duplicate row — the behaviour re-ingestion and ReindexStaleAsync have always assumed. Earlier versions had no such key and appended a second row instead.
Two consequences when pointing the store at a table created by one of those versions:
InitializeAsyncfails fast ifrag_chunksalready contains duplicate(document_id, chunk_index)pairs, because the unique key cannot be created over them. It deletes nothing — the exception carries the duplicate count and the query to inspect them. Decide which row of each pair to keep, remove the rest, then initialize again.StoreAsyncthrows anInvalidOperationExceptionif the table has no unique index on exactly those two columns (theON CONFLICTclause has nothing to infer). Since first use creates that key — and migrates an older table that lacks it — the reachable cause is now narrow: the index was dropped, or the table replaced, after this store had already initialised. A table that simply predates the key is migrated rather than rejected.
Collection names
CreateCollectionAsync caps collection names at 47 characters, shorter than PostgreSQL's own 63-byte identifier limit. Three index names are derived from a collection name — idx_{name}_document_id, idx_{name}_doc_chunk, idx_{name}_embedding — and PostgreSQL silently truncates over-long identifiers rather than failing. 47 is the longest name that leaves all three intact.
The damage does not arrive all at once. Just past the cap only the longest decoration is truncated, so that index merely exists under a name you never asked for. From about 58 characters all three collapse onto the same 63 bytes, and creation then fails partway through: the btree is created first and takes the truncated name, and the unique key is attempted next under what is now the same identifier. That failure is loud — the key's existence check resolves by shape against pg_index, not by name, so it correctly reports no key, then finds the btree squatting on the name and throws a relation named 'X' already exists but is not a unique index. Creation aborts there, before the ANN index is attempted. Rejecting up front is still worth it: that runtime exception names a truncated identifier you never typed and tells you to drop it, when the real fix is a shorter name.
The cap covers only the names the store derives. The plain
CREATE INDEX IF NOT EXISTSstatements —idx_rag_chunks_document_id,idx_rag_chunks_embedding,idx_rag_chunks_sparse— still match on name alone. A hand-made relation already holding one of those names makes the corresponding statement do nothing and the store quietly runs without that index. Unlike the unique key, that costs performance, not correctness, which is why it is documented rather than probed.
Dense index and search behaviour
InitializeAsync builds an HNSW index on embedding with vector_cosine_ops, matching the <=> operator SearchAsync orders by. Two things follow that are worth knowing before tuning anything:
- The index is skipped above 2000 dimensions. pgvector refuses to build HNSW on a wider column (measured against pgvector 0.8.2: 2000 builds, 2001 fails), and
text-embedding-3-largeat 3072 is over the line. Rather than failing initialization for those models, the store leaves the index out and dense search stays an exact sequential scan — slower on large tables, but exact, and the behaviour such deployments already had. To get an index at those widths, storehalfvecor reduce the embedding dimension at the provider. - Where the index exists, dense search is approximate, and results may differ from the exact scan.
A filtered search may return fewer results than TopK, including none. MinScore and MetadataFilter are applied after the HNSW index has picked its candidates, so a selective filter can discard all of them (measured: 2,000 rows, 5 matching the filter, 0 returned). The store mitigates this by setting hnsw.iterative_scan = relaxed_order per query, which makes pgvector keep scanning until it has enough surviving rows (all 5 in that same measurement).
relaxed_orderrelaxes the ordering, not just the scan — that is what it trades for the recall, so pgvector may hand back the surviving rows slightly out of distance order. You never see it: bothSearchAsyncandSearchSparseAsyncread through one method that sorts by score descending before returning, so results are in descending score order under every iterative-scan mode. That matters beyond tidiness — the RRF merge behind ensemble retrieval ranks hits by their position in the list.- Raising
hnsw.ef_searchdoes not compensate — atef_search = 1000the same query still returned 0. Do not reach for it. - Even with iterative scanning, pgvector stops at
hnsw.max_scan_tuples(20,000 by default) rather than degrading into a full scan, so a very selective filter over a very large table can still come back short. Raise that setting if you need it to dig deeper. hnsw.iterative_scanarrived in pgvector 0.8. The store readspg_extension.extversiononce, caches the answer, and simply does not issue the setting against an older extension — so on 0.7 and below a filtered search over an HNSW index keeps the truncating behaviour above. Upgrading pgvector is the only fix.
Similarity score
Dense search returns 1 - (embedding <=> query), which is cosine similarity in [0, 1]. MinScore is applied as a WHERE clause filter before ORDER BY and LIMIT — but where the HNSW index exists that filtering happens after the index picked its candidates, with the consequences described above.
MinScore means something different on the sparse path. SearchSparseAsync scores by raw dot product of matching term weights — unbounded above and not comparable to a cosine similarity. The same option name carries two scales; tune it per path rather than reusing one threshold.
Hybrid search
PgVectorStore does not implement IHybridSearchable. When UseHybridSearch = true, the pipeline falls back to the in-memory BM25 index + RRF merge. See Retrieval — Hybrid search.
Sparse vectors (SPLADE)
Pass enableSparseVectors: true to UsePgVector to register PgVectorSparseVectorStore — a subtype that adds a nullable sparse_embedding sparsevec(N) column to the same rag_chunks rows that hold the dense vectors and serves ISparseSearchable. The dense-only PgVectorStore deliberately does not implement ISparseSearchable, so the pipelines' capability probe is honest and no SPLADE encoding work happens against a store that cannot persist it (the same type-split as Qdrant and Pinecone). Pair it with UseSpladeEncoder — see Sparse retrieval (SPLADE) for the full setup.
services.AddRagNet(rag => rag
.UsePgVector(
connectionString: "Host=localhost;Database=ragdb;Username=postgres;Password=secret",
vectorDimensions: 1536,
enableSparseVectors: true,
sparseVocabularySize: PgVectorSparseVectorStore.DefaultSparseVocabularySize)); // 30522
InitializeAsync then does everything the dense one does, plus:
ALTER TABLE rag_chunks ADD COLUMN IF NOT EXISTS sparse_embedding sparsevec(<vocabularySize>);
CREATE INDEX IF NOT EXISTS idx_rag_chunks_sparse ON rag_chunks USING hnsw (sparse_embedding sparsevec_ip_ops);
Search is server-side: SearchSparseAsync orders by the <#> inner-product operator over that index and returns the negated result as the score. Chunks sharing no term with the query are absent from the results rather than present with score 0.
Requires pgvector 0.7.0 or later — the release that introduced the sparsevec type. InitializeAsync verifies the installed version and throws naming both versions and the upgrade path (ALTER EXTENSION vector UPDATE) rather than letting PostgreSQL's context-free "type sparsevec does not exist" surface later. Note that the iterative-scan mitigation applies to sparse search too and needs 0.8; the sparse HNSW index otherwise truncates filtered queries exactly as the dense one does.
CreateCollectionAsync is sparse-blind
The sparse store inherits CreateCollectionAsync unchanged, and the table it creates has no sparse_embedding column and no sparse index — nothing sparse can be written to or read from a collection. This is a documented sharp edge, not an oversight: collections on this store are already disconnected from the read/write path, because StoreSparseAsync and SearchSparseAsync hardcode rag_chunks exactly as the inherited dense StoreAsync and SearchAsync do. ICollectionManageable here manages tables the store itself never queries. If you need a second sparse-capable table, point a second store at a separate database or schema rather than at a collection.
No ordering contract
Dense and sparse writes for the same chunk can happen in either order. sparse_embedding lives in its own column and is deliberately excluded from the dense upsert's DO UPDATE SET list, so calling StoreAsync after StoreSparseAsync does not drop the chunk's sparse vector. This is the opposite of Pinecone, where an upsert replaces the whole record and the write order is a documented hazard — do not assume that hazard generalises across stores.
Term budget: TopTerms must be ≤ 1000
pgvector's HNSW index rejects a sparsevec with more than 1000 non-zero elements (an unindexed sparsevec column allows up to 16,000). OnnxSpladeOptions.TopTerms defaults to 256, comfortably inside the cap, but raising it past 1000 makes every write fail. StoreSparseAsync validates the whole batch before opening a connection and throws an ArgumentException naming the chunk, the term count and OnnxSpladeOptions.TopTerms — nothing is written. Either lower the term budget, or drop the idx_rag_chunks_sparse index and accept a sequential scan up to 16,000 terms.
sparseVocabularySize is baked in at first initialize
The sparsevec column's dimension is the sparse encoder's vocabulary size, which must be strictly greater than every term id the encoder emits. It defaults to PgVectorSparseVectorStore.DefaultSparseVocabularySize (30522 — BERT's WordPiece vocabulary, which every SPLADE checkpoint in common use inherits).
It cannot be changed in place. ALTER TABLE ... ADD COLUMN IF NOT EXISTS matches on the column name only, so against an existing sparsevec(100) it would leave the dimension alone and report success — after which every sparse write and read fails on the mismatch, and the pipeline swallows both (ingestion survives a sparse-side fault; the ensemble degrades to dense results). The visible outcome would be successful ingestion, successful retrieval, and permanently dense-only quality behind one log line. InitializeAsync therefore fails fast on a dimension mismatch, naming both dimensions and the two ways out:
- construct the store with
sparseVocabularySize:matching the existing column, or ALTER TABLE rag_chunks DROP COLUMN sparse_embeddingand initialize again — which discards every sparse vector already stored. Re-ingest the affected documents to rebuild them;SparseEmbeddingBehaviorrecomputes a sparse vector per chunk on every ingest. (ReindexStaleAsyncalso regenerates sparse vectors from the stored chunk text, but only for documents whose embedding version stamp is stale — dropping the column does not make anything stale, so it will not pick them up on its own.)
So switching to an encoder with a different vocabulary is a drop-and-re-ingest operation, not a config change.
One failure mode has no gate: a term id at or above the declared dimension is rejected by PostgreSQL with a raw
ERROR: sparsevec index out of bounds, naming neither the column nor the option. Nothing guards it because a SPLADE encoder emits ids below its own vocabulary size by construction — if you see this error,sparseVocabularySizeis smaller than your encoder's vocabulary.
Qdrant
Package: Rag.NET.VectorStores.Qdrant
Stores chunks as Qdrant points with a payload. Metadata is stored both as a serialised JSON string in metadata (the round-trip authority) and as individual typed meta_{key} payload fields — numbers as payload doubles, booleans as payload booleans — to enable Qdrant's native payload filtering.
Setup
services.AddRagNet(rag => rag
.UseQdrant(
host: "localhost",
port: 6334,
collectionName: "my-collection",
vectorDimensions: 1536));
Collection initialisation
Call InitializeAsync before first use. It creates the collection with cosine distance if it does not already exist:
var store = provider.GetRequiredService<IVectorStore>() as QdrantVectorStore;
await store!.InitializeAsync();
Metadata filtering
Qdrant filters on meta_{key} payload fields using must conditions matched to the value's kind: strings and dates use keyword match, booleans a boolean match, and numbers a closed gte = lte range (Qdrant's match condition has no double form):
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance", // keyword match on meta_department
["page"] = 3, // numeric range 3 <= meta_page <= 3
},
});
Hybrid search
QdrantVectorStore does not implement IHybridSearchable. When UseHybridSearch = true, the pipeline falls back to the in-memory BM25 index + RRF merge.
Sparse vectors (SPLADE)
Pass enableSparseVectors: true to UseQdrant to register QdrantSparseVectorStore — a subtype that creates the collection with a named sparse vector ("splade") next to the dense vector and serves ISparseSearchable. The dense-only QdrantVectorStore deliberately does not implement ISparseSearchable, so the pipelines' capability probe is honest and no SPLADE encoding work happens against a store that cannot persist it. Sparse vectors live on the same points as the dense embeddings: point ids become deterministic per (DocumentId, ChunkIndex) (making chunk upserts idempotent), and StoreSparseAsync attaches sparse vectors to points previously upserted by StoreAsync — ingestion always calls them in that order. InitializeAsync fails fast when an existing collection was created without sparse support — delete the collection and re-ingest to enable it. See Sparse retrieval (SPLADE) for the full setup including UseSpladeEncoder.
Azure AI Search
Package: Rag.NET.VectorStores.AzureAISearch
Stores chunks as Azure AI Search documents. Implements both IVectorStore and IHybridSearchable — native hybrid search combining BM25 full-text search with HNSW vector search at the service level (as does Weaviate).
Setup
using Azure;
using Rag.NET.AzureAISearch;
services.AddRagNet(rag => rag
.UseAzureAISearch(
endpoint: new Uri("https://my-search.search.windows.net"),
indexName: "my-rag-index",
credential: new AzureKeyCredential("your-api-key"),
vectorDimensions: 1536));
Vector recall: KNearestNeighborsCount
The vector arm's k — how many nearest neighbours Azure retrieves before ranking — is configurable
and defaults to unset:
services.AddRagNet(rag => rag
.UseAzureAISearch(
endpoint: new Uri("https://my-search.search.windows.net"),
indexName: "my-rag-index",
credential: new AzureKeyCredential("your-api-key"),
vectorDimensions: 1536,
configure: o => o.KNearestNeighborsCount = 50));
Unset is a value, not an absence. Microsoft's Create a Vector
Query states that
"both k and top are optional. When unspecified, the default number of results in a response is
50." Omitting the parameter therefore asks for 50 and leaves the number where the platform can
change it.
Before v1.0 this store sent k = TopK, which was worse than sending nothing: at a typical
top-5 it narrowed vector recall to a tenth of Azure's own default, starving RRF fusion of
candidates to fuse and starving any reranker that followed. It now sends nothing unless you ask
(#328).
Set it to 50, or leave it unset, if you enable semantic ranking. The
same page is explicit: "Whenever you use semantic ranking with vectors, set k to 50. Semantic
ranker uses up to 50 matches as input. Specifying less than 50 deprives the semantic ranking
models of necessary inputs." Enabling EnableSemanticRanking together with KNearestNeighborsCount
set below 50 is rejected at UseAzureAISearch registration — ArgumentOutOfRangeException, naming
both — rather than merely discouraged: the two are deliberate settings on the same options object,
made at registration by the same person, and without the guard the failure mode is silent (worse
ranking, no error). Leaving KNearestNeighborsCount unset stays valid and is the right default,
since omitting it is what makes Azure apply its own 50.
Note the asymmetry with TopK: Microsoft documents k as governing "results for vector-only
queries" and top as governing "results for hybrid queries that include a search parameter", so
on the hybrid path this widens the candidate set fusion draws from rather than the number of
results returned.
Index schema
InitializeAsync creates or updates the index with these fields:
| Field | Type | Role |
|---|---|---|
id | String (key) | UUID per chunk |
document_id | String (filterable) | For delete-by-document |
chunk_index | Int32 | Chunk ordinal |
text | SearchableString | Full-text search |
metadata | String | Legacy serialised JSON (still written; read fallback) |
metadata_entries | Collection(Edm.ComplexType) | Typed metadata: {key, stringValue, numberValue, boolValue, dateValue} rows, all sub-fields filterable |
embedding | Collection(Single) | HNSW vector field |
The vector field is configured with an HNSW algorithm profile named "default-algorithm". metadata_entries carries one row per metadata key with the sub-field matching the value's kind populated — which is what makes every metadata key filterable with its type and no per-key schema change: a new key needs no index update, the writer writes it and the filter finds it. Sub-fields of a complex collection cannot be marked sortable (they are multi-valued per document), so sorting by a metadata key is not available.
var store = provider.GetRequiredService<ICollectionManageable>() as AzureAISearchVectorStore;
await store!.InitializeAsync();
Semantic ranking
Opt-in, per store instance, and off by default (#328):
services.AddRagNet(rag => rag
.UseAzureAISearch(
endpoint: new Uri("https://my-search.search.windows.net"),
indexName: "my-rag-index",
credential: new AzureKeyCredential("your-api-key"),
vectorDimensions: 1536,
configure: o => o.EnableSemanticRanking = true));
Enabling it changes the native hybrid path only. SearchAsync is untouched and keeps its
genuine cosine similarity — not as a design preference, but because semantic ranking needs query
text and IVectorStore.SearchAsync has none to give: it takes a ReadOnlyMemory<float> and a
SearchOptions of TopK, MinScore and MetadataFilter. Microsoft is explicit that the ranker
has nothing to work with in that case — "A query with search=* or an empty search string … won't
work because there's nothing to measure semantic relevance against" — so the dense path cannot rank
on any tier in any region (#539).
HybridSearchAsync already carries a textQuery, which is why it is the path that ranks.
With it on:
- The index gains one
SemanticSearchconfiguration (named internally, not a caller-visible knob) prioritising thetextfield as content, built byInitializeAsync/CreateOrUpdateIndexAsyncthe same way every other field is — no separate migration step. - The hybrid query sets
QueryType = SearchQueryType.Semanticand points it at that configuration, so Azure reranks its own fused BM25+vector result. SearchResult.Scorebecomes Azure'sRerankerScoreon that path, returned unrescaled — a roughly 0–4 ordinal relevance score, not a cosine similarity. No new scale declaration was needed:IHybridSearchable.HybridScoreScaleis alreadyScoreScale.OpaqueRanking, because a fused rank and a reranker score are ordinal for the same reason — see Native hybrid scores are ordinal, not similarities.SearchOptions.MinScoreis therefore not applied on the hybrid path, ranked or not: thresholding an ordinal rank as if it were a similarity would keep or drop results arbitrarily.KNearestNeighborsCountmust benull(Azure's own default of 50) or at least 50 — see Vector recall above. Enabling the ranker with a smallerkis rejected at registration rather than left to degrade the ranking silently. The guard still applies: the hybrid query's vector arm has the same recall problem.- A service that cannot rank throws, rather than degrading silently. Azure answers HTTP 200
with ordinary results and no
RerankerScorewhenever the tier is below Basic, the region does not support semantic ranking, or the configuration name does not match — there is no error to catch, only a missing field on results that otherwise look normal.HybridSearchAsynctreats that absence as the error it is:InvalidOperationException, naming the index, rather than handing back plausible-looking results the caller would believe were reranked. This guard is what caught the path mistake above: it fired against a real Azure resource within hours of the feature shipping, on a path that could never have ranked. - A query that cannot reach the native path is refused, not quietly downgraded. The pipeline
dispatches natively only when
UseHybridSearchis set, noEnsembleOptionsis supplied,MinScoreis0.0, and no sparse (SPLADE) arm would run — see How the hybrid path is selected. Violate one and the query would be fused client-side, returning correct results with the ranking silently absent. Because the store declaresIHybridSearchable.NativeOnlyCapability,EnsembleBehaviorthrows instead, naming which of those four settings blocked it.
Resilience and semantic ranking now compose, and did not before #544.
ResilientVectorStoredid not forwardIHybridSearchable, and the pipeline probes the registered (decorated) store, so every hybrid query fell back to client-side fusion — correct results, no ranking, no error. If you read this page before that fix and avoided the combination, you no longer need to:ConfigureResiliencenow yields aResilientHybridVectorStorethat forwards the whole interface and retries the native query like any other read. One caveat: a store implementing bothISparseSearchableandIHybridSearchableis refused at registration with aNotSupportedException, because no decorator variant preserves that pair and silently picking one would re-open the same hole. No shipped store is both.
Native hybrid search
When UseHybridSearch = true, AzureAISearchVectorStore is registered, and the call configures nothing the backend cannot express — no sparse (SPLADE) arm would run, no EnsembleOptions, MinScore left at 0.0 — the pipeline calls HybridSearchAsync. This issues a single Azure AI Search request with both a full-text query and a vectorised query, letting the service perform BM25+vector fusion (the dispatch rule in full):
var results = await pipeline.RetrieveAsync("ISO 27001 audit requirements", new RetrievalOptions
{
TopK = 10,
UseHybridSearch = true,
});
The returned scores are Azure AI Search's own hybrid fusion values — Reciprocal Rank Fusion: each fused query contributes at most about 1/60, so a two-arm hybrid score tops out around 0.033. Not cosine similarities, and exactly why a configured MinScore keeps the client-side path: a similarity-tuned threshold applied store-side to RRF values would silently return nothing. HybridSearchAsync does not forward MinScore to the backend at all, so a caller invoking it directly cannot tune a threshold against the RRF scale either — the fused score is never filtered on this path (see Native hybrid scores are ordinal, not similarities). This is the hybrid path only — plain SearchAsync issues a pure vector query whose score is a bounded function of the similarity metric, which is why the store is treated as similarity-scaled (see Score scale).
Metadata filtering
Filters run against the typed metadata_entries complex collection, one any() clause per filter pair, AND-composed, with the comparison against the sub-field matching the value's kind:
// Generates: metadata_entries/any(m: m/key eq 'department' and m/stringValue eq 'finance')
// and metadata_entries/any(m: m/key eq 'page' and m/numberValue eq 3)
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance",
["page"] = 3,
},
});
This replaces the previous search.ismatch substring probe over the JSON blob with real typed comparisons.
Migrating an existing index
InitializeAsync uses CreateOrUpdateIndexAsync, and adding metadata_entries is an additive field change, so an existing index picks the new field up on the next run without a rebuild. The legacy metadata string field stays declared and written — Azure AI Search forbids removing or re-typing an existing field, and keeping it is what lets the update succeed against a pre-existing index.
Two honest caveats for documents ingested before the change:
- Reading them keeps working. They have no
metadata_entriesrows, so the store falls back to the legacy JSON blob; every value comes back as a string-kindMetadataValue. Nothing throws, nothing is lost. - Filtering them stops working until re-ingest.
MetadataFilternow runs againstmetadata_entries, and old documents have no rows there to match — under the previoussearch.ismatchfiltering they did match. Re-ingest the affected documents to make them filterable again.
Indexing latency
Azure AI Search indexing is near real-time. StoreAsync includes a 1-second delay after batch upload to allow the index to become consistent before a subsequent SearchAsync call. This delay is intentional and sourced from the implementation; plan for it in integration tests.
Weaviate
Package: Rag.NET.VectorStores.Weaviate
Stores chunks as objects of a single Weaviate class (vectorizer: none — Rag.NET brings its own vectors; cosine distance). Implements IVectorStore, IHybridSearchable (native BM25+vector fusion), and ICollectionManageable, all served by one singleton. Object ids are deterministic per (DocumentId, ChunkIndex), so re-ingesting a chunk replaces it.
Local quickstart
docker run -p 8080:8080 \
-e AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED=true \
-e PERSISTENCE_DATA_PATH=/var/lib/weaviate \
-e DEFAULT_VECTORIZER_MODULE=none \
cr.weaviate.io/semitechnologies/weaviate:latest
Setup
services.AddRagNet(rag => rag
.UseWeaviate(
endpoint: new Uri("http://localhost:8080"),
className: "RagChunks", // capital letter + letters/digits/underscores
vectorDimensions: 1536));
The class name doubles as a GraphQL field, so Weaviate requires a capitalized GraphQL-valid name (validated eagerly at registration). Optional settings via the configure callback:
services.AddRagNet(rag => rag
.UseWeaviate(new Uri("https://my-cluster.weaviate.cloud"), "RagChunks", 1536, options =>
{
options.ApiKey = "wcs-api-key"; // sent as Authorization: Bearer
options.Tenant = "customer_a"; // opt into multi-tenancy
}));
Class schema and initialisation
InitializeAsync creates the class if missing: fixed properties document_id (text, field tokenization so Equal filters match whole ids), chunk_index (int), text (text — feeds BM25), and metadata_json (serialised metadata for lossless round-tripping). StoreAsync also initialises lazily on first write, so a forgotten InitializeAsync can never let Weaviate's auto-schema create the class with the wrong tokenization.
var store = provider.GetRequiredService<IVectorStore>() as WeaviateVectorStore;
await store!.InitializeAsync();
Scores
Dense search maps Weaviate's cosine distance (0 = identical … 2 = opposite) to Score = 1 - distance / 2, so an identical vector scores 1.0. Hybrid search returns Weaviate's relative-score-fusion value, already in [0, 1] but ordinal rather than a similarity. MinScore is applied to the dense path's mapped score only — the hybrid path's fused score is never thresholded by it (see Native hybrid scores are ordinal, not similarities).
Native hybrid search
When UseHybridSearch = true and the call configures nothing the backend cannot express — no sparse (SPLADE) arm would run, no EnsembleOptions, MinScore left at 0.0 — the pipeline calls HybridSearchAsync: a single GraphQL hybrid: {query, vector} request lets Weaviate fuse BM25 and vector rankings server-side, so a chunk that matches only by keyword is still found. The returned scores are Weaviate's relative-score-fusion values in [0, 1], not cosine similarities. See Retrieval — How the hybrid path is selected for the dispatch rule in full.
Metadata filtering and auto-schema
Each chunk metadata key is written as an extra typed meta_{key} property — text for strings, number for numbers, boolean for booleans; Weaviate's auto-schema (enabled by default in the official image) adds these properties with the matching type on first write, making them server-side filterable with typed where operands (valueText, valueNumber, valueBoolean). The metadata_json property remains the lossless round-trip authority:
// Generates: where: {path: ["meta_department"], operator: Equal, valueText: "finance"}
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance",
},
});
Multiple filter entries are wrapped in a single And operand. Note that auto-schema created meta_* properties use Weaviate's default word tokenization, so Equal on a multi-word value matches per token — keep filterable metadata values single-token.
Multi-tenancy
Set WeaviateOptions.Tenant to isolate data per tenant: the class is created with multiTenancyConfig: {enabled: true}, the tenant itself is created during initialisation (idempotent), and every store/search/delete carries it. Two stores configured with different tenants on the same class never see each other's chunks.
Chroma
Package: Rag.NET.VectorStores.Chroma
Stores chunks as records of a single Chroma collection via the REST v2 API — deliberately the lightweight, dense-only adapter. Implements IVectorStore and ICollectionManageable, served by one singleton. Record ids are {documentId}:{chunkIndex}, so re-ingesting a chunk upserts (replaces) it; the chunk text rides as the record's document and metadata carries the chunk metadata plus document_id and chunk_index.
Local quickstart
docker run -p 8000:8000 chromadb/chroma
Setup
services.AddRagNet(rag => rag
.UseChroma(
endpoint: new Uri("http://localhost:8000"),
collectionName: "rag-chunks")); // 3-512 chars: letters/digits/._-, alphanumeric ends
The collection is created automatically (with the cosine space) on first use; Chroma infers vector dimensions from the first upsert, so no dimension parameter is needed. Optional settings via the configure callback:
services.AddRagNet(rag => rag
.UseChroma(new Uri("http://localhost:8000"), "rag-chunks", options =>
{
options.Tenant = "my_tenant"; // default: default_tenant
options.Database = "my_database"; // default: default_database
options.ApiKey = "static-token"; // sent as Authorization: Bearer
}));
Chroma addresses collections by UUID internally; the store resolves the configured name to its UUID once and caches it. If the collection is deleted or recreated behind the store's back, the next operation transparently re-resolves and retries once.
Scores
Chroma returns cosine distance = 1 - cosine similarity (0 = identical … 2 = opposite), mapped to Score = 1 - distance, so an identical vector scores 1.0 and an orthogonal one 0.0 (opposite vectors go negative). MinScore is applied to the converted score.
The store requires the cosine space. If the configured collection already exists with a different space (Chroma's default is squared L2), the first operation fails fast with an InvalidOperationException naming the actual space — the score conversion would otherwise be silently on the wrong scale and MinScore would misfilter. Delete and recreate the collection (re-ingesting its documents) or point the store at a cosine collection.
Hybrid search
ChromaVectorStore does not implement IHybridSearchable (or ISparseSearchable) — Chroma has no native BM25+vector fusion for externally supplied embeddings. When UseHybridSearch = true, the pipeline falls back to the in-memory BM25 index + RRF merge; if you want native hybrid or sparse search, use Qdrant, Pinecone, or PgVector (sparse/SPLADE), Weaviate, or Azure AI Search instead. See Retrieval — Hybrid search.
Metadata filtering
Chunk metadata keys are stored as-is on each record with native value types (numbers as numbers, booleans as booleans; dates as a $date:-prefixed sentinel string, since record values cannot be objects) and filtered server-side with Chroma's typed $eq operator; multiple filter entries are composed with $and:
// Generates: where: {"$and": [{"department": {"$eq": "finance"}}, {"team": {"$eq": "core"}}]}
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance",
["team"] = "core",
},
});
Note that document_id and chunk_index are reserved record-metadata keys (a same-named chunk metadata key would be overwritten by them).
Pinecone
Package: Rag.NET.VectorStores.Pinecone
Stores chunks as records of a Pinecone serverless index via the official Pinecone.Client SDK. Implements IVectorStore and ICollectionManageable, served by one singleton; the opt-in sparse variant adds ISparseSearchable (see below). Record ids are {documentId}:{chunkIndex}, so re-ingesting a chunk upserts (replaces) it.
Pinecone stores no document body: the chunk text lives in record metadata (key text) next to document_id and chunk_index, and is read back into SearchResult.Chunk.Text. Keep chunks comfortably under Pinecone's ~40 KB metadata limit per record — text plus all metadata must fit.
SDK version note: the package pins
Pinecone.Client3.1.0, not 4.x. The 4.x control-plane models require avector_typeresponse field that Pinecone Local does not send, so index create/describe/list fail against the emulator (pinecone-dotnet-client#54; the SDK repository was archived in July 2026, so no fix is expected).This pin was verified empirically against Pinecone Local only — the store's behaviour on the live Pinecone service has not been exercised in this repository's test suite. Be aware that 3.1.0 sends
X-Pinecone-API-Version: 2025-01, which is past Pinecone's 12-month guaranteed support window for an API version. Upgrading to 4.x (a source-breaking change: renamed metric enums and reshaped describe/list responses) becomes necessary if Pinecone retires 2025-01 or if you need 2025-04+ features — native sparse indexes in particular — and would cost Pinecone Local coverage for index management until #54 is resolved.
Setup
services.AddRagNet(rag => rag
.UsePinecone(
apiKey: "your-api-key",
indexName: "rag-chunks", // 1-45 chars: lowercase letters/digits/-, alphanumeric ends
vectorDimensions: 1536));
Optional settings via the configure callback:
services.AddRagNet(rag => rag
.UsePinecone("your-api-key", "rag-chunks", 1536, options =>
{
options.Namespace = "customer-a"; // namespace isolation (see below)
options.EnableSparseVectors = true; // sparse variant — dotproduct index required
options.Cloud = ServerlessSpecCloud.Aws; // serverless placement, default aws
options.Region = "us-east-1"; // ... default us-east-1
options.Endpoint = new Uri("http://localhost:5080"); // Pinecone Local
}));
Index lifecycle
CreateCollectionAsync(name, dimensions) creates a serverless index (cloud/region from the options; cosine metric — dotproduct when sparse vectors are enabled) and polls describe until the index reports ready, bounded by PineconeOptions.IndexReadyTimeout (default 2 minutes; serverless creation typically takes under a minute, Pinecone Local is ready instantly). Deleting a missing index is a no-op; storing or searching against a missing index fails fast with an exception naming CreateCollectionAsync as the fix.
Local development (Pinecone Local)
Pinecone Local is an in-memory emulator of the control and data planes — no account or API key needed (keys are accepted and ignored):
docker run -p 5080-5090:5080-5090 \
-e PORT=5080 -e PINECONE_HOST=localhost \
ghcr.io/pinecone-io/pinecone-local:latest
Point the store at it with options.Endpoint = new Uri("http://localhost:5080") — the http scheme also switches the SDK's data-plane gRPC channels to plaintext. The emulator serves the control plane on port 5080 and gives every index its own data-plane port from 5081–5090, advertised as localhost:{port} — hence the port-range publish (and a cap of ten live indexes). Emulator limitations to plan around: data is not persisted across restarts, at most 100,000 records per index, and no sparse values on dense indexes (see the sparse section below); delete-by-metadata-filter is rejected exactly like the real serverless service.
Scores
Pinecone returns native similarity scores, so MinScore applies directly: cosine similarity in [-1, 1] on the default metric (identical vector ⇒ 1.0, orthogonal ⇒ 0.0). On a dotproduct index (sparse variant) dense scores are raw dot products and sparse scores are sums of matching term-weight products — both unbounded above, so tune MinScore for that scale.
Namespace isolation
Set PineconeOptions.Namespace to scope every upsert, query, and delete to one Pinecone namespace — the features.md "namespace-based collection isolation". Two stores configured with different namespaces on the same index never see each other's chunks, and DeleteByDocumentIdAsync only deletes within its own namespace. Leave it null for the default namespace.
Metadata filtering
Chunk metadata keys are stored as-is on each record with native value types (numbers as numbers, booleans as booleans; dates as a $date:-prefixed sentinel string, since record values cannot be objects) and filtered server-side with Pinecone's typed $eq operator; multiple filter entries are composed with $and:
// Generates: filter: {"$and": [{"department": {"$eq": "finance"}}, {"team": {"$eq": "core"}}]}
var results = await pipeline.RetrieveAsync("query", new RetrievalOptions
{
MetadataFilter = new Dictionary<string, MetadataValue>
{
["department"] = "finance",
["team"] = "core",
},
});
document_id, chunk_index, and text are reserved record-metadata keys (a same-named chunk metadata key would be overwritten by them).
Delete by document
Serverless indexes do not support delete-by-metadata-filter (the service answers "Serverless and Starter indexes do not support deleting with metadata filtering" — Pinecone Local included), so DeleteByDocumentIdAsync lists vector ids by the {documentId}: prefix and deletes by id, in batches. Ids whose remainder after the prefix is not purely digits are skipped — they belong to a longer document id that merely starts the same way (e.g. deleting doc never touches doc:7's chunks).
Sparse vectors (SPLADE)
Set EnableSparseVectors = true in the configure callback to register PineconeSparseVectorStore — a subtype that serves ISparseSearchable next to the dense interfaces. The dense-only PineconeVectorStore deliberately does not implement ISparseSearchable, so the pipelines' capability probe is honest and no SPLADE encoding work happens against a store that cannot persist it (the same type-split as Qdrant and PgVector). Pair it with UseSpladeEncoder — see Sparse retrieval (SPLADE) for the full setup.
Sparse values ride on the same records as the dense embeddings: StoreSparseAsync upserts the full record (dense + sparse + metadata), so it needs no prior StoreAsync call for that chunk. Order matters in the other direction, though: a Pinecone upsert replaces the entire record and StoreAsync writes records without sparse values, so calling StoreAsync after StoreSparseAsync for the same chunk silently drops its sparse vector. Always store dense first, sparse second — Rag.NET's own ingestion (StorageBehavior) and RegenerateSparseAsync both do, so this only bites hand-rolled write paths; re-ingesting a chunk means re-running both steps. SearchSparseAsync issues a sparse query with an all-zero dense vector sized from the live index (Pinecone requires a dense vector on every query; zeroing it nulls the dense contribution — the documented alpha = 0 weighting).
Pinecone only accepts sparse values on dotproduct indexes. The sparse variant's CreateCollectionAsync therefore creates dotproduct indexes, and its first data-plane use fails fast with an InvalidOperationException naming the fix when the configured index has a different metric — the real service would accept sparse upserts into a cosine index and only reject at query time.
Pinecone Local gap: the emulator rejects sparse values on writes to dense indexes (its gRPC upsert answers INVALID_ARGUMENT; its REST path silently drops them), though it does serve sparse queries. Concretely, the container suite covers: dotproduct index creation by the sparse variant, dense store/search through it, the non-dotproduct fail-fast, and sparse querying (including that the zero dense vector is sized from the live index rather than the configured VectorDimensions). It does not cover storing sparse values or the sparse store-then-search round-trip — that test is skipped with this reason, so the same-record sparse write path has only been verified by construction, never executed against a live Pinecone serverless index. Treat it as unproven until you run it against the real service.
Redis (RediSearch)
Package: Rag.NET.VectorStores.Redis
Stores chunks as Redis hashes under {index}:{documentId}:{chunkIndex} and queries them with an HNSW vector index. The point is reuse: a large share of .NET applications already run Redis for caching, and for those teams pointing it at retrieval is a much lower bar than standing up a second datastore — the same argument that justifies pgvector.
Needs the RediSearch module: Redis Stack, or Redis 8 and later where it is built in. Plain Redis answers FT.CREATE with an unknown-command error.
Setup
services.AddRagNet(rag => rag
.UseRedis(
configuration: "localhost:6379",
indexName: "ragnet-idx",
vectorDimensions: 1536));
If Redis is already in the application — the case this store exists for — hand it the connection you have. The store does not dispose a multiplexer it did not create:
var redis = ConnectionMultiplexer.Connect("localhost:6379");
services.AddRagNet(rag => rag.UseRedis(redis, "ragnet-idx", 1536));
Call InitializeAsync before first use. It creates the index only if absent, because re-creating it would discard every stored vector.
Similarity score
RediSearch returns vector_score as a cosine distance in [0, 2] — 0 is identical, larger is worse. That is the opposite direction from every score in this library, so the store converts it to 1 - distance and reports ordinary cosine similarity in [-1, 1].
This is why RedisVectorStore does not implement IScoreScaleAware: its scores are already on the scale every threshold assumes, exactly as pgvector's 1 - (embedding <=> $1) is. Publishing the distance unconverted would invert every ranking and silently break MinScore — an integration test asserts an exact match scores ~1.0 rather than ~0.0 for precisely that reason.
Hybrid search is declined, not approximated
RediSearch can run a text query alongside the vector one, but its text scoring is TF-IDF-shaped rather than the BM25 the hybrid arm fuses. A store advertising IHybridSearchable here would be fusing a score it cannot describe, so this one does not — and the pipeline falls back to its own BM25 arm, which is honest about what it is.
Metadata: persisted, and filterable on declared keys
Chunk metadata is persisted as a metadata hash field (a JSON blob) and comes back from both the keyed lookup and search. document_id is also indexed as its own TAG (which is how DeleteByDocumentIdAsync finds a document's chunks without scanning keys the store did not write) and chunk_index as NUMERIC.
Filtering on arbitrary metadata keys, however, needs those keys declared when the store is constructed:
services.AddRagNet(rag => rag
.UseRedis(
configuration: "localhost:6379",
indexName: "ragnet-idx",
vectorDimensions: 1536,
filterableMetadataKeys: ["tenant"]));
Each declared key becomes a case-sensitive md_<key> TAG attribute, and SearchAsync applies MetadataFilter server-side, inside the KNN query — not as a post-filter on the top-K page. This is Redis's one real limitation here, and it is a hard one: RediSearch matches only against attributes its schema names, so a filter naming a key that was not declared throws InvalidOperationException rather than silently returning an unfiltered page, and an index built before a key was declared fails InitializeAsync — it must be recreated and re-ingested once the key is added. See the package README for the full explanation of why Redis alone needs this.
Multi-index federation
Package: Rag.NET (core)
FederatedVectorStore wraps two or more IVectorStore instances behind a single store, so you can search across collections living in different backends (e.g. a private PgVector index plus a shared Qdrant index) without migrating data. It is registered as the IVectorStore, so the entire pipeline (MMR, reranking, caching, …) composes unchanged.
Setup
services.AddRagNet(rag => rag
.UseFederatedSearch(f => f
.AddStore(_ => new PgVectorStore("Host=...;Database=private", 1536), "private-pg")
.AddStore(_ => new QdrantVectorStore("localhost", 6334, "shared", 1536), "shared-qdrant")
.WithPrimary(0) // optional: writes/deletes target this store (default: first)
.WithRrfK(60))); // optional: RRF constant (default: 60)
At least two stores are required (validated at registration). Store factories receive the IServiceProvider and run once, when the federated store is first resolved.
Behaviour
- Search fans out to all stores concurrently, then merges the per-store rankings with N-way Reciprocal Rank Fusion: each hit contributes
1 / (k + rank)(1-based rank,k=RrfK) and the mergedScoreis the summed RRF score, not a cosine similarity.TopKis applied after the merge. Ties on the merged score are broken deterministically: the chunk that first appeared in the lower store index wins, then the lower per-store rank. MinScoreis applied by each store against its own similarity scale before fusion; the mergedScoreis RRF. Beware cross-backend coherence: the sameMinScorevalue means different things to different backends (e.g. raw cosine similarity in[0, 1]for PgVector/Qdrant vs. Azure AI Search's rescaled relevance score, which for cosine bottoms out around0.333), so a threshold tuned for one store may over- or under-filter another. This is the store-sideSearchOptions.MinScoreapplied during fan-out, which is unaffected by the score scale the federated store declares for its own merged results.- Provenance: every merged result's chunk metadata gains a
source.storeentry with the store's name (fromAddStore(..., name)) or its zero-based index. The source store's own chunk is never mutated — the tag is written into a copied metadata dictionary. - Writes and deletes go to the primary store only.
DeleteByDocumentIdAsyncdoes not touch secondary stores — documents ingested directly into secondaries must be deleted there. - Degraded, never broken: a store that throws during search is skipped with a logged warning; the federated search itself only throws (
InvalidOperationExceptionnaming the stores) when every store failed.
Interaction with other registrations
UseFederatedSearch supersedes any earlier IVectorStore registration (standard last-wins container semantics). Do not combine it with UsePgVector/UseQdrant-style calls — add those stores through the builder instead.
Persistent conversation memory: UsePersistentMemory resolves the DI IVectorStore and normally filters recalled exchanges by PersistentMemoryOptions.MinScore (default 0.7), a threshold calibrated to the similarity scale. Federated results carry RRF scores (about 0.033 at best for two stores), so that threshold would discard every match. FederatedVectorStore therefore declares IScoreScaleAware with ScoreScale.OpaqueRanking, and persistent memory reacts by skipping MinScore entirely: it injects the store's top TopK matches in rank order and logs one warning per memory instance naming the store type and the ignored threshold. Recall works against a federated store; what you give up is the ability to require a minimum relevance — every recall injects the best TopK the federation returns, however weak. Lower TopK (default 3) if that is too much context, or point persistent memory at a dedicated similarity-scaled store when you need a real threshold.
Score scale (IScoreScaleAware)
IScoreScaleAware is an opt-in capability interface declaring what a store's SearchResult.Score means:
| Scale | Meaning | Declared by |
|---|---|---|
ScoreScale.Similarity | Comparable, roughly [0, 1], safe to threshold against a fixed cut-off | The assumed default — every store that does not implement the interface, AzureAISearchVectorStore included |
ScoreScale.OpaqueRanking | Ordinal only; magnitude is not comparable and must not be thresholded | FederatedVectorStore (RRF sums) |
The declaration describes IVectorStore.SearchAsync — the interface IScoreScaleAware sits on — not any capability method the store also happens to offer. Consumers probe with store is IScoreScaleAware { ScoreScale: ScoreScale.OpaqueRanking }. Every other store in the library, AzureAISearchVectorStore included, is treated as similarity-scaled, so SearchOptions.MinScore on the retrieval path behaves exactly as before; the probe currently affects persistent conversation memory only. IHybridSearchable carries the sibling declaration for the hybrid path: HybridScoreScale defaults to the same ScoreScale.OpaqueRanking, and Azure AI Search and Weaviate both take that default for HybridSearchAsync — see Native hybrid scores are ordinal, not similarities.
Azure AI Search does not implement IScoreScaleAware, and that is the correct declaration
rather than an omission. Since #539 moved
semantic ranking to HybridSearchAsync, its SearchAsync returns a genuine cosine similarity under
every option — so the value it would declare is ScoreScale.Similarity, which is exactly what this
interface's absence already means. The store implemented it while the ranker sat on the dense path
and the value depended on a constructor argument; with that gone, implementing it would leave a
member indistinguishable from not implementing it. The hybrid declaration is the one that still
carries information, and it stays: IHybridSearchable.HybridScoreScale is ScoreScale.OpaqueRanking,
covering the fused score and the reranker score alike.
SearchAsync issues a pure vector query under every option, whose @search.score is a bounded
monotone function of the similarity metric and is thresholdable, which is why Similarity — reached
through the interface's absence — is the correct answer for it. Its hybrid scores
(HybridSearchAsync) are a different matter — RRF values around 1/60 per fused query, or Azure's
0–4 RerankerScore when semantic ranking is on — and are covered by
IHybridSearchable.HybridScoreScale instead, which is OpaqueRanking for both; the pipeline's hybrid dispatch never
applies a non-zero MinScore to them (a configured MinScore keeps the client-side path) and
neither does the store itself when called directly — see Native hybrid scores are ordinal, not
similarities.
Limitations
Federation is dense-only in this release: IHybridSearchable (native hybrid), sparse search, and ICollectionManageable capabilities of the underlying stores are not federated. When UseHybridSearch = true, the pipeline's BM25 fallback still applies over the shared in-memory/SQLite BM25 index, not per federated store.
Implementing a custom vector store
See Extending for the full guide.