Proto commits in izihawa/summa

These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)

Commit:08d79d8
Author:Pasha Podolsky
Committer:GitHub

Report summa-server readiness, commit on shutdown, accept common_grams in SDL, serve partial partition reads; prepare 2.0.2 (#207) * Report summa-server readiness and commit admitted documents on shutdown summa-server opened indexes lazily, so a restarted shard accepted TCP connections minutes before it could answer: requests waited for the first open while Kubernetes and the broker treated it as ready. It now opens every index at startup and serves grpc.health.v1, NOT_SERVING until opening finishes and again once shutdown begins. Graceful shutdown discarded the uncommitted generation, so writers that commit on a schedule lost everything admitted since their last commit on each restart. Shutdown now commits each index before stopping its writer. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Accept common_grams in SDL Common word pairs could only be enabled through SchemaBuilder, so indexes created from SDL (CreateIndex) could not use them. Text fields now accept indexed<token_position, common_grams: ["..."]>, and GetIndexInfo renders the list so the schema round-trips. Admission and the segment builder keep validating the field and the words. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Serve partial reads of partitioned indexes when partitions are down A partitioned read failed whenever any partition lacked a routable replica, so restarting one shard took the whole index offline. With --partial-partition-reads the broker serves Search and GetTextStats from the partitions that are up: partitions without a routable replica or answering UNAVAILABLE are left out of both the statistics and the search, reported in the new missing_partitions response field and counted in summa_broker_partial_reads_total. Other failures and losing every partition still fail; GetIndexInfo, writes and commits stay strict, and GetDocument reports UNAVAILABLE rather than NOT_FOUND when the document may be on a missing partition. The default stays strict. The real-server tests waited for an empty index list before creating indexes, which the broker answers before discovering its backends; they now wait until every backend is healthy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prepare version 2.0.2 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Commit:fb87fcf
Author:pasha
Committer:pasha

Serve partial reads of partitioned indexes when partitions are down A partitioned read failed whenever any partition lacked a routable replica, so restarting one shard took the whole index offline. With --partial-partition-reads the broker serves Search and GetTextStats from the partitions that are up: partitions without a routable replica or answering UNAVAILABLE are left out of both the statistics and the search, reported in the new missing_partitions response field and counted in summa_broker_partial_reads_total. Other failures and losing every partition still fail; GetIndexInfo, writes and commits stay strict, and GetDocument reports UNAVAILABLE rather than NOT_FOUND when the document may be on a missing partition. The default stays strict. The real-server tests waited for an empty index list before creating indexes, which the broker answers before discovering its backends; they now wait until every backend is healthy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Commit:84390de
Author:pasha

Rebrand search engine as Summa and restore its documentation site

Commit:976f951
Author:pasha

Merge main and preserve shared readers across compact storage layouts

Commit:37fa66c
Author:pasha

Add compact Seismic search and single-copy binary vector storage

Commit:2fa28f8
Author:pasha

Fix large singleton upserts and share payloads across deletion reloads

Commit:b3d1979
Author:pasha

fix(search): seed real passages for document-only candidates

Commit:d2ac2e8
Author:pasha

Add row deletion, upserts, and opt-in compaction

Commit:c45944f
Author:pasha

Add symbolic L1 formulas, RRF attribution and search tracing

Commit:22487da
Author:pasha

Complete L1 scoring with optional BMP forward storage

Commit:974d83f
Author:pasha
Committer:pasha

Coordinate global fusion and reapply shared linear scoring across shards

Commit:4c3bd07
Author:pasha
Committer:pasha

Add cross-vertical candidate backfill and request-driven linear L1 ranking

Commit:da98358
Author:pasha

fix: unify text pruning and stream ordered phrase results

Commit:e74f888
Author:Pasha Podolsky
Committer:GitHub

feat(tokenizer): ICU segmentation, light stemmers, same-position variants (#158) * feat(tokenizer): ICU segmentation, light stemmers, same-position variants - segmenter: icu (ICU4X word segmenter: dictionary words for Chinese and Japanese with their bigrams as variants, LSTM breaks for Thai, Lao, Khmer, Burmese), NFKC + lowercase, Arabic normalisation, ё -> е - stem: none | light | snowball; light = ports of Lucene's light/minimal stemmers (en, fr, de, es, it, pt, ru, fi, hu, sv, no, ar) - keep_original: the written word is the token, its stem and folded form are variants at the same position (Token::variant, excluded from field length); fold: diacritic folding as variant or in place - max_token_length (default 64): long tokens dropped, position kept - Tokenizer::tokenize_query(text, hint, exact): match queries use the stem, phrases and terms the written form; server and wasm converters pick by clause kind - roadmap tokenization section rewritten * feat(tokenizer): t2s option folds traditional Chinese to simplified (OpenCC table) * feat(tokenizer): morph option: Japanese and Korean dictionary morphology (lindera) UniDic and ko-dic embedded behind the cjk-dict feature (hermes-server builds with it). Hangul runs and kana runs (Han when hinted ja) are analysed into morphemes: particles, auxiliaries, endings and suffixes are dropped with their positions kept; content morphemes are indexed as written with the base form and bigrams as same-position variants; match queries use the base form, phrases the surface. A morph spec fails to parse without the feature. * build(server): ca-certificates for lindera dictionary downloads * refactor(tokenizer): lex(...) spec, one LexOptions, one tokenize_with API

Commit:7acaa5b
Author:Pasha Podolsky
Committer:GitHub

feat: lexical vertical: stop words, positions v2, windowed MaxScore, searcher-wide DF (#156) * feat(tokenizer): stop_words spec with gap-preserving phrase offsets - TokenizerSpec gains stop_words: true|false; DynamicStemmer drops the routed language's stop words before stemming without renumbering the surviving tokens, so phrase distances survive (quantum@0 art@3). - PhraseQuery carries per-term offsets (with_offsets); server, query language and WASM pass tokenizer positions through; the scorer checks start + delta instead of +1 per term. - Boolean conversion drops a phrase clause (or a Boolean made only of such clauses) that tokenizes to nothing, instead of failing the request. - Tests: tokenizer gaps/routing/spec round-trip, index-level phrase gaps and field lengths, converter offsets and dropped clauses. * docs: lexical vertical roadmap (positions v2, pruning, field-level reorder, tokenization) * feat(core): cursor-addressed position stream (positions v2) - New per-term position stream: sorted per-document positions delta-coded and packed in 128-value rounded-width blocks, block offset table, footer with total count and magic. No per-term skip list, no ordinal bits needed for chunked fields. - BlockPostingList gains an extended footer (total_positions, flags, magic) and, for terms with positions, a u64 cursor per L0 block (values before the block); BlockPostingIterator tracks the in-block tf prefix and exposes position_cursor(). Legacy footers without the magic still parse. - Readers open positions zero-copy (TermPositions: Stream | Legacy) and decode only the one or two blocks covering [cursor, cursor + tf) per candidate; PhraseScorer and TermScorer address positions by cursor. - Merges re-pack every source into one stream and rebase cursors by the values of preceding sources; legacy positioned sources are decoded and re-encoded so merged terms always address a v2 stream. - Tests: stream round trip across blocks, re-pack equivalence, legacy bridge, cursor persistence through serialize/seek/streaming merge, legacy footer parse and concat, mixed-format refusal; existing phrase, chunked and merge suites pass unchanged. Synthetic mid-frequency term (20k chunks, tf 1-4 over ~200 tokens): legacy 112112 B vs v2 53032 B + 1256 B cursors (2.25 vs 1.09 B/position). * docs: reorder rewrites the position stream per chunk run * feat(core): persisted field lengths and length-aware block bounds for BM25 - .chunks version 2: per-field sections of kind 1 hold one u16 token count per document of the segment for every plain indexed text field, written by the builder (multi-valued fields sum their values) and concatenated on merge with zero fill for sources without the column. Version 1 files still read. - Plain fields now score with their real length in MaxScore, TermScorer and PhraseScorer (LengthSource / DocLengths); tf-as-length remains only for legacy segments without norms. - BlockPostingList: the fourth L0 word packs (max_tf u16, min_len u16) and the footer carries the list minimum (FLAG_LEN_BOUNDS); bounds are bm25(max_tf, idf, min_len, avg) instead of the 1-b floor. Builder derives min_len from chunk lengths or norms, merges copy the words and take the minimum, lists rebuilt without lengths store 1, legacy f32 words convert on merge. A cursor uses the stored minimum only when it scores with real lengths, so the bound always dominates the score. - Tests: chunk map v2 round trip, zero-fill merge, v1 read; packed bounds through serialize and streaming merge; index-level brute-force BM25 parity with rank-safe MaxScore over 600 mixed-length docs; norms across a merge. * feat(tokenizer): unicode segmenter (UAX #29, CJK bigrams, folding) and NLTK stop lists - stem(..., segmenter: unicode): UAX #29 word boundaries (float-zero -> float, zero; p53, co2, 10.1007 stay whole), character bigrams over runs of Han, Hiragana and Katakana with one position each, and NFKD diacritic folding of Latin, Cyrillic and Greek tokens applied after stemming so stop lists and Snowball see the original letters. simple stays the default so existing fields keep their tokenization; queries use the same segmenter. - Stop lists: enable the stop-words crate's nltk feature. Its default ISO English list (~1300 words) removes content words such as "state"; the NLTK lists (150-250 words) cover all 18 routed languages, including Tamil, which no longer falls back to English. - Tests: word-boundary splitting and offsets, folding after stemming, Cyrillic folding, stop-word gaps in unicode mode, CJK bigram runs and phrase positions, spec parse/render round trip. * feat(query): filters and phrases as predicates of text MaxScore - Boolean planner: when every SHOULD clause is a text term and the MUST/MUST_NOT clauses combine into one document bitset, the text MaxScore executors (plain and chunked, one group per field) run with the bitset as their predicate. The top-k is exact over the filtered documents; before, text SHOULD went through an over-fetched unfiltered top-k thinned by a PredicatedScorer, which could drop phrase matches ranked below the candidate budget. - Documents matching only the MUST clauses fill the tail with score 0 (BitsetFillScorer) when fewer than limit scored documents survive, so Boolean semantics are unchanged. - PhraseQuery::as_doc_bitset / bitset_cardinality_estimate: a quoted span becomes a bitset (chunk ids resolved to documents); chunked phrase scorers keep every matching document instead of their top limit, which a MUST constraint requires. DocBitset::next_set_bit for ordered iteration. - Tests: chunked field with phrase + fast-field filter + OR-of-phrases wrapper and ordinals; plain field with phrase + filter and score-0 fill; existing MUST+SHOULD expectations unchanged. * feat(query): score phrases by phrase frequency PhraseScorer counts the occurrences of the whole phrase in the unit and scores BM25 over that count with the summed idf of the terms and the unit's real length (Lucene semantics), replacing 1.5 x BM25 over the summed term frequencies. Test: phrase_scores_by_phrase_frequency (lands with the next commit's test file). * fix(query): rank-safe block skips in Block-Max MaxScore The block-max branch moved every cursor at the minimum document to its next block whenever their block bounds could not reach the threshold, even when another essential cursor still held a document inside that block. That document was later scored without the skipped cursors' contributions and could drop out of the top-k. The skip is now bounded by the next essential document (the Block-Max MaxScore rule): a cursor skips its block only when the block ends before that document and otherwise seeks to it inside the block. The loop is shared by text and sparse MaxScore. Regression test block_max_skip_never_jumps_over_another_essential_cursor reproduces the drop (top hit missing before the fix) against brute-force BM25; also lands phrase_scores_by_phrase_frequency. * feat(core): superblock bounds for text MaxScore - Every L1 group of eight blocks stores a packed (max_tf, min_len) word after the L1 last_doc entries (FLAG_L1_BOUNDS, 4 B per 1,024 postings), derived from the packed L0 words at build and on merge. Legacy lists have none. - MaxScore: when the block bounds of the cursors at the minimum document fail the threshold, their group bounds are checked too; if those fail as well, a cursor jumps to its next group instead of its next block, still bounded by the next essential document (rank-safe). Sparse cursors keep single-block skips. groups_skipped is reported next to blocks_skipped. - The brute-force BM25 parity test now runs on 4,000 documents so block and superblock skips fire. * feat(schema): per-field BM25 parameters - SDL: indexed<k1: .., b: ..> on text fields (b validated to 0..=1, rejected on non-text fields); FieldEntry.bm25_k1 / bm25_b (serde optional, metadata unchanged when absent); SchemaBuilder::set_bm25_params; GetIndexInfo renders them back. - query::Bm25Params (replaces the unused duplicate in traits) carries k1/b through the MaxScore cursors and executors (scores and block/group bounds), TermScorer and PhraseScorer; Bm25Params::for_field resolves a field's values with the Lucene defaults. Stored block bounds are parameter-free, so k1/b change without a rebuild. - Tests: SDL parse and validation, SDL round trip through the server renderer, index-level b = 0 (length-free scores through term and MaxScore paths) and k1 = 0.5 score check. * feat(query): sequential-dependence proximity rescoring - MatchQuery.proximity_weight / proximity_window (default window 8): BooleanQuery::with_proximity carries the config into the text MaxScore finishers, which over-fetch 4x the limit, add a bonus per adjacent query term pair from ordered (adjacent) and unordered (within the window, half weight) position windows, saturated with the field's k1/b and length and weighted by the pair's mean idf, then re-sort and cut to the limit. Positions are read through the term cursors from the v2 stream. Works on plain and chunked fields and under filters; cross-segment threshold seeding is disabled for the pass because the bonus lifts scores above the BM25 floor. Without the sync feature the stage is a no-op. - Tests: window counting, config defaults, index-level ordering for plain and chunked fields (adjacent > windowed > distant; ties without the stage; limit honoured after rescoring; with a MUST filter), converter round trip of the proto fields. * feat(core): field-level BP reordering of chunked text fields A chunked text field with the reorder attribute gets its own Recursive Graph Bisection order over its virtual chunk ids (segment/text_reorder.rs), driven by the standalone reorder pass (optimizer, IndexWriter::reorder): - forward index from the field's own postings (terms with df >= 2, highest df dropped to the memory budget), ForwardIndex::from_csr, existing BP with the block size as partition; - the term dictionary, postings and positions are rewritten term by term: the planned field's postings sorted by new virtual id with cursors and bounds from the permuted lengths, its positions re-encoded in that order, every other field's terms copied through the merger's single-source path; - .chunks written with the planned field's rows in the new order, other chunk maps and length columns unchanged. Document ids, the store, the fast fields and every vector field do not move; the segment's converged flag combines the sparse and text passes. Test: 600 docs in two interleaved topical clusters; identical documents, scores and (for exact-match queries) ordinals before and after the pass, the chunk map leaves indexing order, and a merge after the pass still resolves phrases and filters. * feat(query): approximate, capped, and budgeted text MaxScore - MatchQuery.heap_factor: approximate text top-k via threshold scaling; an approximate pass neither seeds from nor publishes the shared floor - MatchQuery.max_terms: keep the highest-idf terms of a long match - SearchRequest.time_budget_ms / SearchResponse.truncated: anytime mode; the deadline travels on SharedThreshold, text executors check it every 4096 iterations and return the best-so-far - Searcher: search_with_{positions,count}_budgeted - roadmap step 6 documented * feat(query): searcher-wide and cross-shard BM25 statistics - Query::text_terms lists the BM25 terms of a query; a multi-segment searcher sums their document frequencies over its segments through the term dictionary and scores every segment with the same IDF, corpus size and average length (ScorerOptions::global_stats) - GetTextStats RPC + SearchRequest.text_stats: the cross-shard statistics contract for a scatter-gather broker; the broker proxies both today - roadmap: Lucene/turbopuffer parity checklist * perf(query): windowed Block-Max MaxScore for text cursors Window-at-a-time execution (Lucene MaxScoreBulkScorer, turbopuffer batched iterator advancement): per-window block-max partition from skip entries, window skips without decode, essential lists bulk-scored into a dense buffer + bitset, branch-free competitive filtering, non-essential lists sought to survivors one iterator at a time. Rank-safe; parity test against an exhaustive scorer and the document-at-a-time loop. Sparse cursors keep the loop. Roadmap: Lucene/turbopuffer parity checklist filled in. * feat(query): query term de-duplication with query-tf weights Repeated match tokens collapse into one clause boosted by their count; TermQueryInfo.weight scales a boosted term's idf so it stays on the text MaxScore path with the same scores as the repeated clauses. * build: gate text reorder to native, regenerate client proto stubs * build(ts-client): carry the new MatchQuery and SearchRequest fields

Commit:d487b77
Author:Pasha Podolsky
Committer:GitHub

feat: dynamic per-document stemmer and wire-level PhraseQuery (#152) Text fields can declare text<stem(by: <field>, default: <language|simple>)>: the segment builder reads all values of the hint field from each document and stems that document's tokens per script (Snowball stemmers are script-local, so Cyrillic+Latin documents stem both parts); TermQuery/MatchQuery gain a tokenizer_hint the server passes to the same tokenizer so query and index stemming agree. Unknown tokenizer names and dangling hint fields now fail schema parsing instead of silently falling back to lowercasing ("default" is registered as the documented alias of "simple"). PhraseQuery is exposed on the wire (oneof field 12): server-side tokenization with the field's tokenizer, consecutive terms with optional slop, BM25 scored, degrading to a MUST of terms on fields without positions. The query-language parser builds the same PhraseQuery for quoted spans instead of an AND. Clients (Python, TypeScript, WASM) accept phrase and tokenizer_hint; generated bindings regenerated. Design: docs/dynamic-tokenizer-and-phrase.md.

Commit:1e38dd0
Author:Pasha Podolsky
Committer:GitHub

feat(vector): add streaming ScaNN indexes

Commit:fa97e03
Author:Pasha Podolsky
Committer:GitHub

feat(broker): hermes-broker — route the hermes.proto surface across instances (#139) * feat(broker): add hermes-broker routing the hermes.proto surface across instances A stateless gRPC broker that fronts many hermes-server backends behind one address. Serves the exact hermes.proto SearchService/IndexService (clients switch by re-pointing endpoints) plus a broker-only hermes.broker admin proto and grpc.health.v1. - Discovery: kubernetes pod watch (shard id + role pod labels) or static --backend list; index->backend routing learned via ListIndexes polling with a Suspect/Evicted health machine (grace, two-probe recovery). - Routing: per-index pass-through; writes to the shard master only; CreateIndex placed by glob rules (dated families follow their shard); ambiguous multi-shard names read deterministically and refuse writes until a rule pins them. - Contract: byte-faithful write responses, grpc-timeout propagated with no broker-imposed deadlines, admission mirrors --max-concurrent-searches with the server's exact RESOURCE_EXHAUSTED message, ListIndexes served from cache. - Tests: unit (topology/placement/health/deadline math), integration (real broker binary vs scripted mock backends: pass-through equality, stream re-routing, eviction/recovery, deadline headers), env-gated e2e against real hermes-server subprocesses (CI runs it after building the server). - CI publishes ghcr.io/spacefrontiers/hermes/hermes-broker alongside the server image; both Dockerfiles now copy the new workspace member. * feat(broker): add gRPC server reflection grpcurl-driven verification (the deployment runbook's parity checks) needs reflection: neither the broker nor hermes-server previously answered schema-less tooling. Registers both the hermes.proto and hermes-broker.proto descriptor sets. * fix(broker): install rustls crypto provider; exit on discovery crash kube's rustls-tls stack panics at first TLS use unless a process-level CryptoProvider is installed (rustls 0.23); install ring explicitly. A discovery-task failure (error or panic) now drains and exits the server via the internal shutdown channel instead of leaving a zombie broker answering NOT_SERVING forever.

Commit:8b6a471
Author:Pasha Podolsky
Committer:GitHub

fix(core): make LogSumExp combiner a count-invariant smooth maximum (#123) The raw log-sum-exp added ln(n)/t per document, so chunk COUNT outranked chunk quality: a 300-chunk compendium of ~0.5 matches scored 0.55 + ln(300)/1.5 ~= 4.3 and buried every focused paper (0.72 + ln(3)/1.5 ~= 1.4). In production this made "off-label aripiprazole usage" return generic long documents with zero aripiprazole results in the hybrid fusion. LogSumExp now computes the softmax-weighted average sum(softmax(t*s)*s): - n identical scores combine to that score (count-invariant) - always bounded by the max - still tracks a dominant score at any scale (the naive LSE - ln(n)/t would collapse a single 15-score sparse chunk among 299 zeros to ~6.4; the softmax form keeps it at ~14.9) Verified live on documents_20260726: COMBINER_MAX (the t->inf limit of the new form) returns the aripiprazole papers at 0.70-0.72 where the old default returned 280-470k-char compendia. Re-pinned two tests that had enshrined the count boost as their observable: SOAR dedup is now asserted through exact-score equality and TQ document completeness through Sum. Score scale for multi-value fields changes (bounded by max); downstream absolute-score thresholds should be re-checked.

Commit:d0280c8
Author:pasha

Separate heap, mapped, and pinned memory stats

Commit:2001cad
Author:pasha

Optimize BMP compression and LSP search

Commit:1d15c5b
Author:pasha
Committer:pasha

fix(search): unify candidate budgets and BMP metadata

Commit:bd33302
Author:Pasha Podolsky
Committer:GitHub

perf(search): unify candidate overfetch budget (#92)

Commit:0ea0cdf
Author:Pasha Podolsky
Committer:GitHub

feat(vector): make ANN retraining atomic and scalable (#89)

Commit:6e38f42
Author:Pasha Podolsky
Committer:GitHub

feat: chunk-level fusion with cross-vertical corroboration (#16) * perf: query-path madvise discipline, SOAR wiring, hybrid RRF fusion - MADV_RANDOM on scattered-access mmap regions (BMP block data/doc maps, ANN-backed flat vectors, set once at segment open) and batched MADV_WILLNEED prefetch of rerank candidates in both rerank paths - restore query-pattern advice on BMP merge sources after merge reads - wire SOAR into SDL (soar: selective|full|aggressive|off), IVF trainers, and search-side dedup of spilled candidates (in-place, allocation-free) - add union score fusion: FusionMethod (RRF / normalized weighted sum) and Searcher::search_fused; share RRF formula with the L2 reranker * feat: expose hybrid fusion through gRPC, WASM, and client SDKs - proto: FusionQuery (weighted sub-queries, RRF / normalized weighted sum, rrf_k, fetch_limit) as a top-level-only Query variant - server: fusion branch in SearchService::search — fans out sub-queries, fuses ranked lists (union), composes with the L2 reranker (fused list becomes the L1 pool); nested fusion rejected in convert_query - wasm: structured search accepts { fusion: { queries, method, rrfK, fetchLimit } } with offset support - clients: fusion query builders for Python and TypeScript; regenerated proto stubs * feat: doc_mass sparse cropping, extended RaBitQ, binary IVF, sync binary scorer - sparse: doc_mass document-side mass cropping (keep top-weight entries covering the configured fraction of each vector's mass); avg sparse vector length in hermes-tool info and GetIndexInfo vector_stats - dense: extended multi-bit RaBitQ (bits: 2-8 in SDL) — magnitude refinement codes with a full-precision-query estimator; <5% mean relative distance error at 5 bits vs ~2x more at 1 bit - binary: fix sync scorer (production INVALID_ARGUMENT on multi-thread runtimes) with multi-thread regression test; new BinaryIvfIndex — k-majority Hamming IVF with exact in-cluster codes, SDL indexed<ivf>, built at commit past build_threshold, rebuilt on merge, rayon-parallel assignment - docs: unified-vector-quantization design note; CLAUDE.md development and testing rules * feat: chunk-level fusion with cross-vertical corroboration Fusion now runs at (doc, ordinal) chunk granularity: per-chunk scores are collected from each sub-query, ranked by chunk score within each vertical, RRF-fused per chunk key, and combined into doc scores with a MultiValueCombiner (default Max). - fixes the junk-vertical regression: a doc rank-1 in sparse is no longer outvoted by a mediocre doc present in both lists on different chunks (pinned by test) - same-chunk hits across verticals compound (ordinal-aligned schemas) - fused results carry per-chunk ordinal_scores again (chunk attribution for snippets survived doc-level fusion loss) - default per-sub-query fetch depth raised to max(4x limit, 50) - proto: FusionQuery.combiner (unset = MAX); clients updated

Commit:fd4ec44
Author:Pasha Podolsky
Committer:GitHub

feat: doc_mass cropping, extended RaBitQ, binary IVF, sync binary scorer fix (#15) * perf: query-path madvise discipline, SOAR wiring, hybrid RRF fusion - MADV_RANDOM on scattered-access mmap regions (BMP block data/doc maps, ANN-backed flat vectors, set once at segment open) and batched MADV_WILLNEED prefetch of rerank candidates in both rerank paths - restore query-pattern advice on BMP merge sources after merge reads - wire SOAR into SDL (soar: selective|full|aggressive|off), IVF trainers, and search-side dedup of spilled candidates (in-place, allocation-free) - add union score fusion: FusionMethod (RRF / normalized weighted sum) and Searcher::search_fused; share RRF formula with the L2 reranker * feat: expose hybrid fusion through gRPC, WASM, and client SDKs - proto: FusionQuery (weighted sub-queries, RRF / normalized weighted sum, rrf_k, fetch_limit) as a top-level-only Query variant - server: fusion branch in SearchService::search — fans out sub-queries, fuses ranked lists (union), composes with the L2 reranker (fused list becomes the L1 pool); nested fusion rejected in convert_query - wasm: structured search accepts { fusion: { queries, method, rrfK, fetchLimit } } with offset support - clients: fusion query builders for Python and TypeScript; regenerated proto stubs * feat: doc_mass sparse cropping, extended RaBitQ, binary IVF, sync binary scorer - sparse: doc_mass document-side mass cropping (keep top-weight entries covering the configured fraction of each vector's mass); avg sparse vector length in hermes-tool info and GetIndexInfo vector_stats - dense: extended multi-bit RaBitQ (bits: 2-8 in SDL) — magnitude refinement codes with a full-precision-query estimator; <5% mean relative distance error at 5 bits vs ~2x more at 1 bit - binary: fix sync scorer (production INVALID_ARGUMENT on multi-thread runtimes) with multi-thread regression test; new BinaryIvfIndex — k-majority Hamming IVF with exact in-cluster codes, SDL indexed<ivf>, built at commit past build_threshold, rebuilt on merge, rayon-parallel assignment - docs: unified-vector-quantization design note; CLAUDE.md development and testing rules

Commit:6fcd904
Author:Pasha Podolsky
Committer:GitHub

feat: expose hybrid fusion through gRPC, WASM, and client SDKs (#14) * perf: query-path madvise discipline, SOAR wiring, hybrid RRF fusion - MADV_RANDOM on scattered-access mmap regions (BMP block data/doc maps, ANN-backed flat vectors, set once at segment open) and batched MADV_WILLNEED prefetch of rerank candidates in both rerank paths - restore query-pattern advice on BMP merge sources after merge reads - wire SOAR into SDL (soar: selective|full|aggressive|off), IVF trainers, and search-side dedup of spilled candidates (in-place, allocation-free) - add union score fusion: FusionMethod (RRF / normalized weighted sum) and Searcher::search_fused; share RRF formula with the L2 reranker * feat: expose hybrid fusion through gRPC, WASM, and client SDKs - proto: FusionQuery (weighted sub-queries, RRF / normalized weighted sum, rrf_k, fetch_limit) as a top-level-only Query variant - server: fusion branch in SearchService::search — fans out sub-queries, fuses ranked lists (union), composes with the L2 reranker (fused list becomes the L1 pool); nested fusion rejected in convert_query - wasm: structured search accepts { fusion: { queries, method, rrfK, fetchLimit } } with offset support - clients: fusion query builders for Python and TypeScript; regenerated proto stubs

Commit:ac8b9e7
Author:pasha

feat: add Reciprocal Rank Fusion (RRF) to reranker Fuse L1 (first-stage) and L2 (reranker) rankings using RRF: score(d) = 1/(k + rank_L1) + 1/(k + rank_L2) Enabled via rrf_k field on Reranker config (0 = disabled, typical: 60). Works for both dense and binary vector reranking paths. Also fixes partial_cmp → total_cmp consistency and removes unnecessary clone in binary reranker path.

Commit:673f8e5
Author:pasha
Committer:pasha

feat: add BinaryDenseVector field type with Hamming distance search Native binary dense vector support for compact hash-based similarity: - Packed-bit storage (dim/8 bytes per vector), scored via XOR + popcount - Brute-force flat search with batch Hamming distance scoring - Unified .vectors file format (single TOC+footer for dense + binary) - Shared VectorResultScorer for both dense and binary queries - Unified RerankerConfig dispatching to cosine or Hamming reranking - Full stack: proto, gRPC server, Python client, TypeScript client - SDL syntax: field hash: binary_dense_vector<768> [indexed]

Commit:bb2acd4
Author:pasha
Committer:pasha

feat: add PrefixQuery for wildcard prefix matching Add PrefixQuery that matches all documents containing terms starting with a given prefix (e.g., `site:https://reddit.com/r/Trans*`). Materializes union of matching posting lists into sorted doc ID set with O(log N) seek. Score is 1.0 (filter-style). - SSTable prefix_scan with block-level early termination - SegmentReader get_prefix_postings for field-scoped prefix lookup - PrefixQuery/PrefixScorer with SortedVecDocSet - QL grammar: prefix_query rule with URL-friendly charset, implicit OR - gRPC proto: PrefixQuery message, MatchQuery trailing * promotion - Python and TypeScript client support

Commit:8e76550
Author:pasha
Committer:pasha

feat: record-level BMP reorder command with SimHash clustering Two-phase streaming rebuild that shuffles individual ordinals across blocks so similar vectors (by SimHash) cluster together, enabling better block pruning during search. - reorder_bmp_blob(): Phase 1 computes per-vid SimHash from block postings using weighted per-bit-plane accumulators; Phase 2 writes reordered blob with O(n) per-source-block gather using FxHashMap slot mapping - Exposed via IndexWriter::reorder(), gRPC Reorder RPC, CLI subcommand, and TypeScript/Python clients - SimHash unit tests (identity, similarity, topic clustering, avalanche) and integration test verifying intra-block clustering quality - Benchmark (bmp_reorder) measuring before/after latency on 100K-1M docs - Code quality: extracted stafford_mix() and push_bmp_field_toc() shared helpers, removed unused inv_perm allocation, pre-allocated grid_entries

Commit:7d3e010
Author:pasha
Committer:pasha

perf: SIMD correctness fixes, indexing optimizations, and sparse query improvements - Fix FMA feature gate on squared_euclidean_avx2 (SIGILL crash on non-FMA CPUs) - Fix Bits32 over-read in unpack_rounded_delta_decode - Fix missing returns after NEON dispatch in unpack_32bit and add_one - Add round-to-nearest-even for f32_to_f16 conversion - Multi-accumulator patterns for squared_euclidean (NEON/SSE/AVX2) and fused_dot_norm_f16 - SSE dequantize_uint8: use _mm_cvtepu8_epi32 instead of scalar loads - NEON unpack_8bit_delta_decode: use widening chain instead of scalar loads - Eliminate double hash lookups in segment builder (entry API pattern) - Lower default compression level from 7 to 3 - Fix PK refresh error propagation (try_join_all instead of filter_map) - Fix PK rollback on closed channel - Lazy ordinal decoding in MaxScore executor (defer until scoring phase) - Cache last_doc_id on SparseBlock to avoid redundant decode in serialize()

Commit:64b0f6a
Author:pasha
Committer:pasha

feat: add primary key deduplication for unique document constraints Add `primary` SDL attribute and bloom-filter-backed dedup check on document insertion, rejecting duplicates against both committed segments (via fast-field text dictionary) and uncommitted in-flight documents. - SDL: `primary` attribute (text fields only, at most one per index) - PrimaryKeyIndex: bloom filter + FxHashSet for O(1) negative lookups - Writer integration: check_and_insert in add_document, refresh on commit, clear on abort - Proto: DocumentError for per-document error reporting in batch/stream - Server: partial-success semantics (continue indexing on dup errors) - Python client: return errors from index_documents responses

Commit:a03c24f
Author:pasha
Committer:pasha

fix(proto): support multi-value fields in gRPC responses FieldValueList wrapper replaces bare FieldValue in SearchHit.fields and GetDocumentResponse.fields, preserving all values for multi-value fields. Updated server, Python client, and TypeScript client.

Commit:0779995
Author:pasha
Committer:pasha

query review: bug fixes, perf optimizations, client cleanup Bug fixes: - TermQuery on non-text fields now returns error (server + QL parser) - BooleanQuery single-clause optimization: unwrap single MUST/SHOULD - count_estimate SHOULD sum: use saturating_add to prevent u32 overflow - RangeScorer::seek() backwards after exhaustion: add exhaustion guard - TopKResultScorer::size_hint(): return remaining count, not total Performance: - RangeScorer: cache fast_field reference at construction, eliminates per-doc HashMap lookup in matches() during scan_forward() Verified: 532 tests pass, 4 crates compile clean

Commit:f7974ed
Author:pasha
Committer:pasha

refactor: change rerank_factor from uint32/usize to float/f32 Allow fractional rerank factors (e.g. 1.5x, 2.5x candidates). All fetch_k computations now use ceil(k * rerank_factor) for rounding. Changes: - proto: DenseVectorQuery.rerank_factor uint32 -> float - core: DenseVectorQuery.rerank_factor usize -> f32, default 3.0 - core: segment reader fetch_k = ceil(k * rerank_factor.max(1.0)) - core: QL parser rerank param parsed as f32 - server: converter passes f32 directly (no usize cast) - python: DenseVectorQuery.rerank_factor type int -> float - typescript: regenerated proto stubs (rerankFactor already number)

Commit:486ecf5
Author:pasha
Committer:pasha

fix: prevent silent store block loss, add doc count verification at build and merge time - EagerParallelStoreWriter: propagate compression thread panics via resume_unwind instead of silently dropping blocks (root cause of progressive document loss after merges) - build_store_streaming: verify store doc count matches expected count - merge_store: hard error on per-segment store/meta num_docs mismatch (prevents postings/store doc_id desync in merged output) - merger: verify post-merge store num_docs matches metadata total - Add test_store_fields_survive_multiple_merges (450 docs, 3 rounds) - Add test_store_large_scale_multi_merge (3200 docs, 4 rounds, MmapDir) - Switch test_needle_combined_all_modalities to MmapDirectory

Commit:1530ca1
Author:pasha
Committer:pasha

feat: adaptive block sizes, matryoshka reranking, config tuning - Add DP optimal partition for sparse posting blocks (partitioner.rs) Allowed sizes [16,32,64,128,256], minimizes block-max waste via O(N×5) DP - Wire adaptive partitioning into segment builder (replaces fixed block_size) - Add MAX_BLOCK_SIZE=256, from_postings_with_partition() to block.rs - Bump decode buffer capacities 128→256 in scoring.rs - Add matryoshka_dims pre-filter to L2 dense reranker Scores on truncated dims first, keeps top final_limit×2, then exact cosine Zero-copy SIMD scoring directly from raw_buf (no intermediate buffers) - Expose matryoshka_dims in proto Reranker message and server converter - Server defaults: indexing threads cpu/4, memory 16GB

Commit:f65a34a
Author:pasha
Committer:pasha

perf: sparse vector query-time optimizations - Binary search on skip entries in LazyTermCursor::seek (O(log n) vs O(n)) - Precompute abs_query_weight in LazyTermCursor, SparseTermScorer, BMP - find_partition: binary search via partition_point (O(log n) vs O(n)) - BMP: running total_remaining sum (O(1) vs O(n) per iteration) - BMP: pre-resolve skip_start, skip ordinals decode when ordinal_bits==0 - Threshold-aware non-essential early exit in LazyBlockMaxScore - SparseSkipList::find_block: linear scan -> binary search - SparseIndex::load_block_direct: skip dim_id lookup per block load - BlockSparsePostingIterator: lazy ordinal decode, pre-alloc buffers - BlockSparsePostingIterator: direct indexing in doc()/weight() hot paths - Refactor builder/mod.rs: extract dense/sparse/postings/store build modules

Commit:b1bd22d
Author:pasha
Committer:pasha

feat: add RetrainVectorIndex RPC to proto, server, Python and TypeScript clients - Added RetrainVectorIndex RPC to IndexService in hermes.proto - Implemented server handler calling IndexWriter::rebuild_vector_index() - Added retrain_vector_index() to Python client - Added retrainVectorIndex() to TypeScript client - Regenerated proto stubs for both clients

Commit:49ce22a
Author:pasha
Committer:pasha

feat: add hermes-client-typescript with CI/CD and npm publish support

Commit:3ddbf44
Author:pasha

feat: add ListIndexes RPC to proto, server, and Python client

Commit:843cbb5
Author:pasha
Committer:pasha

feat: add dense vector reranking (L1/L2 two-stage retrieval) Add a Reranker to SearchRequest that enables two-stage retrieval: L1 runs any query (text/sparse/bool) to get candidates, then L2 reranks by exact dense vector distance on stored vectors.

Commit:4f5683f
Author:pasha
Committer:pasha

feat: add stress testing framework and memory instrumentation - Add Python stress test framework (hermes-client-python/stress_test/) - Support for sparse, dense, fulltext, and mixed index types - Memory monitoring via RSS tracking of server process - Metrics collection: throughput, latency percentiles, segment analysis - CLI with configurable parameters - Fix IndexWriter caching in gRPC server - Cache IndexWriter in IndexRegistry to reuse across requests - Previously each batch_index_documents created new writer, causing segment fragmentation (200 segments for 5k docs -> now 4-8 segments) - Add memory breakdown logging on flush - Log postings, sparse vectors, dense vectors, interner, positions - Helps diagnose memory usage by component - Add reload interval configuration - New --reload-interval-ms CLI arg for hermes-server - Configurable via IndexConfig.reload_interval_ms - Add memory stats proto messages (for future GetIndexInfo enhancement) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

Commit:c1416de
Author:pasha
Committer:pasha

Add new combiners

Commit:52d8d7b
Author:pasha
Committer:pasha

Support ordinals

Commit:315228d
Author:pasha
Committer:pasha

dev

Commit:6ea336e
Author:pasha
Committer:pasha

Dev

Commit:0440ba4
Author:pasha
Committer:pasha

Update server

Commit:cc66e99
Author:pasha

Update server

Commit:240b73b
Author:pasha
Committer:pasha

Fix NCCL

Commit:2650d74
Author:pasha
Committer:pasha

Add optimizations

Commit:6ab8cbf
Author:pasha
Committer:pasha

Development

Commit:8180824
Author:pasha
Committer:pasha

Add dense

Commit:197f0c8
Author:pasha
Committer:pasha

Add sparse

Commit:4d58a4c
Author:pasha

init

Commit:d25b6ee
Author:Pasha Podolsky

[dev] Return back autofields and mapped fields

The documentation is generated from this commit.

Commit:dd9b41c
Author:Pasha Podolsky

[dev] Restore WASM crate

Commit:9a26b20
Author:Pasha Podolsky

update tantivy

Commit:50f4e30
Author:Pasha Podolsky
Committer:Pasha Podolsky

dev

Commit:e2e1e8c
Author:Pasha Podolsky

[feat] Add skip_updated_at_modification for indexing

Commit:0437a34
Author:Pasha Perevedentsev

feat: Auto ID fields

Commit:2670257
Author:pasha

rename removed_fields

Commit:f5a44a6
Author:pasha

feat: Develop constraints

Commit:e3cc49a
Author:pasha

feat: Add public API prototype, add grpc-web

Commit:e9bba6e
Author:pasha

feat: Update tantivy

Commit:807d3c4
Author:pasha
Committer:pasha

feat: Cache development, part 2

Commit:7823ea0
Author:pasha

feat: Added caching layer for collectors

Commit:6493a4e
Author:pasha

feat: Make hotcache creation optional

Commit:98a736f
Author:pasha

feat: Streaming with queries

Commit:ffec1e4
Author:pasha
Committer:pasha

fix: Deal with lowercase in term query, refactor aggregations

Commit:45c9f7b
Author:pasha

feat: Remove excessive multi-index search requests

Commit:13fce33
Author:pasha

remove ner

Commit:972eec3
Author:pasha
Committer:pasha

feat: Replace missing_field_policy with the list of removed fields

Commit:741249b
Author:pasha
Committer:pasha

fix: Duplicating auto fields

Commit:cf902b6
Author:pasha

feat: Support do nothing policy

Commit:65193d3
Author:pasha

feat: Support dynamic field mapping

Commit:7b144a8
Author:pasha
Committer:pasha

feat: Reduce reliance on attributes, added customazable term_field_mappers

Commit:73c245e
Author:pasha
Committer:pasha

More morphology support

Commit:9e73e0c
Author:pasha
Committer:pasha

feat: Add default for new configs

Commit:164310a
Author:pasha

feat: Refactor configuring query parsers, add initial inflection support

Commit:2008d69
Author:pasha

release: 0.15.19

Commit:efb3902
Author:pasha

feat: Add field aliases

Commit:6822cfe
Author:pasha
Committer:pasha

feat: Update Tantivy, add ipfs-hamt-directory for large directories in IPFS, new release

Commit:2a5b8b5
Author:pasha

feat: Refactor top_docs

Commit:3194304
Author:pasha

fix: More logging and refactor API names

Commit:6bab67b
Author:pasha

feat: Added opstamp to hotcache, updated Tantivy

Commit:815cc72
Author:pasha
Committer:pasha

feat: Bump version

Commit:e0b17a1
Author:pasha

Implements #158. ExistsQuery and regular expressions in grammar

Commit:e05f964
Author:pasha

feat: Excluded segments

Commit:13e07bb
Author:pasha

fix: Align documentation with the latest changes

Commit:fc471a1
Author:pasha

feat: End `iroh` experiment

Commit:623b64a
Author:pasha

feat: A log of bugfixes for WASM-build

Commit:7b1f099
Author:pasha

feat: SummaQL fixes

Commit:dfa2693
Author:pasha

feat: Update to beetle

Commit:7870eb9
Author:pasha
Committer:pasha

feat: Added SummaQL, modify docs

Commit:2f508ca
Author:pasha

feat: Added more outputs to CLI

Commit:aabf4d6
Author:pasha

feat: Move `default_fields` to MatchQuery

Commit:d7ee102
Author:pasha
Committer:pasha

feat: Remove extra channel, support higher compression

Commit:c970e06
Author:pasha
Committer:pasha

feat: Updating Tantivy