OPEN ENGINEERING RESEARCH LAB · HOBBY & LEARNING PROJECT · 1,081 TESTS

Anvi Search: Enterprise Commerce Search & Endeca Migration

Anvi Search is enterprise commerce search built as an Endeca migration instrument — delivering hybrid vector retrieval, a visual merchandiser workbench, and 100% in-VPC data sovereignty.

Engineered for multi-million SKU retailers and B2B distributors modernizing off legacy Oracle Endeca (MDEX) and Lucidworks Fusion (SolrCloud) without rewriting business merchandising rules, losing dynamic cartridges, or surrendering proprietary catalog data to multi-tenant cloud SaaS.

💡

Open Engineering & Learning Laboratory: Anvi Search is an independent engineering research project built as a hobby, technical playground, and deep learning endeavor by a Principal Search Architect. It is not a commercial product for sale. Everything is engineered from first principles and shared openly to explore how modern hybrid retrieval, vector embeddings, and an explainable merchandiser workbench can solve enterprise search modernization.

अन्वी
The Origin & Meaning of "Anvi"ROOT: अन्वेषण (ANVESHANA)

Derived from the ancient Sanskrit root अन्वेषण (Anveshana) / अन्वी (Anvi) — meaning "To Seek, To Inquire, To Follow the Path of Knowledge, and To Explore Truth". In classical Indian thought, Anvi represents one who bridges and guides the way through deep, uncharted territory. Anvi Search was named to honor this pure spirit of discovery: engineering an open, transparent, and explainable search discovery engine from first principles.

THE STRATEGIC CHOICE

The Enterprise Search Dilemma: Three Paths Forward

Thousands of enterprise retailers and B2B distributors run billions in commerce revenue through aging Oracle Endeca (MDEX) and Lucidworks Fusion (SolrCloud) clusters. Here is how staying, moving to SaaS, and Anvi Search compare:

PATH 1: STAY ON ENDECA / FUSIONHIGH RISK

Staying on Legacy Endeca or Lucidworks Fusion

Oracle MDEX is End-of-Life, and Lucidworks Fusion faces enterprise sunsetting and licensing uncertainty. Sustaining support is expensive, ZooKeeper quorum splits and JVM GC freezes cause outages, and finding engineers who know Forge or SolrCloud internals is nearly impossible.

  • ❌ End-of-Life and platform sunsetting support risks
  • ❌ Fragile ZooKeeper quorums, JVM memory bloat & 4-hour batch reindexing
  • ❌ Lacks sub-12ms hybrid vector embeddings and native MCP agent protocols
PATH 2: MOVE TO SAASCOSTLY LOCK-IN

Moving to Cloud SaaS (Coveo/Algolia)

Cloud search vendors promise turnkey AI, but require rewriting all your merchandising rules from scratch. Query logs, customer PII, and proprietary catalog margins leave your private network, and usage-based pricing scales to hundreds of thousands every year.

  • ❌ $100,000 to $300,000+ recurring annual query fees
  • ❌ Proprietary catalog and customer order data leaves your VPC
  • ❌ Black-box ML models you cannot inspect or explain
PATH 3: ANVI SEARCHTHE MODERN STANDARD

Drop-In Parity for Endeca & Lucidworks + Modern Hybrid RAG

Built specifically as a dual migration instrument for Oracle Endeca (MDEX) and Lucidworks Fusion (SolrCloud). Keeps 100% of your business rules, dynamic cartridges, and query pipeline stages. Upgrades you to sub-12ms Hybrid Vector search, inside your own private VPC with zero recurring SaaS fees.

  • ✓ 100% automated rule, cartridge & pipeline stage parity (0 retraining)
  • ✓ 100% In-VPC data sovereignty (zero cloud query fees)
  • ✓ Sub-12ms Hybrid BM25 + Vector RRF with explainable scoring
FULL-STACK ARCHITECTURAL SUPERPOWERS

The 6 Core Superpowers of Anvi Search

Combining cutting-edge NLP, Machine Learning, and Contextual RAG with enterprise-grade Rule Management, Visual CMS Cartridges, and 100% On-Premise / In-VPC sovereignty.

🔮

1. Hybrid Vector RRF & Dynamic Swatches

Dense semantic embeddings meet lexical SKU precision in sub-12ms:

  • Parent-Child SKU Collapse: Matches on child attributes (size, spec, color) while rolling up to 1 clean parent card to eliminate result clutter.
  • Query-Aware Swatch Promotion: Searching "red running shoes" dynamically promotes the red variant image, rather than the default black thumbnail.
  • Application-Layer RRF: Dual-branch BM25 + 384-dim dense vector fusion preserving sort stack determinism (Price, Margin) with neural MiniLM cross-encoder reranking.
🤝

2. Geo-Climate & Order History Personalization

Coveo-class relationship & location boosting with 100% In-VPC privacy:

  • Geo-Climate Location Intent: Searching "jacket" in New York in November automatically boosts heavy down parkas (+300 pts), while in Dallas (72°F) it boosts lightweight fleece and windbreakers.
  • First-Party Order History Affinity: Binds past brand loyalty (e.g. 3 Arc'teryx orders → +250 boost), saved size profiles (Size 10.5 → +150), and price tiers at query time.
  • Installed-Base & Contract Graphs: Connects registered equipment (e.g. CAT 320D) to compatible parts (+500 pts) and B2B contract tiers with zero customer PII leakage to cloud SaaS.

3. Kafka Event Streaming (vs Push/Pull)

Replacing 4-hour batch ITL crawls and table locks with real-time CDC:

  • Recommended Flagship Route: Enterprise ERP/PIM systems emit domain events to partitioned Kafka topics. Zero search cluster lock contention.
  • Sub-20ms CDC Hot Path: Price and inventory changes stream continuously into in-memory document overlays.
  • Pluggable Ingestion Adapters: Supports Debezium DB CDC (Oracle, Postgres, Mongo), REST Bulk Push, and S3/SFTP batch pull feeds.
🌐

4. Multi-Site & Multi-Locale Architecture

Hierarchical rule inheritance across brands, countries, and channels:

  • Inheritance Rule Tree: Define global clearance suppression once at the Global Enterprise level; child country sites inherit automatically.
  • Local Site Overrides: UK or German storefronts can override banners and campaigns without modifying parent rules (depth-before-breadth).
  • Language-Family Cores: Dedicated tokenizers and embeddings for English, German, and Japanese, with site-filtered slices inside each core.
🏬

5. Store-Based BOPIS & Pricing Overlays

Omnichannel retail without the 700 million document explosion trap:

  • In-Memory Store Bitsets: Supports 2,000 physical stores by layering store inventory over a single store-agnostic catalog core.
  • Sub-Millisecond BOPIS Filtering: Shopper clicks "In Stock at Chicago Store #1042" → filter evaluates in 0.4ms with zero index bloat.
  • Store-Specific Clearance Pricing: Overlays local store markdown prices over national online MSRP.
🔒

6. 100% On-Prem & In-VPC (Zero SaaS)

Complete data sovereignty, air-gapped resilience, and $0 query fees:

  • Runs in Your Own Infrastructure: Deploy on AWS, Azure, GCP, or bare-metal Kubernetes. Customer queries and margins never leave your firewall.
  • Zero SaaS Query Taxes: Flat, predictable infrastructure costs. High-traffic flash sales never generate surprise $200k SaaS bills.
  • Air-Gapped & Transparent: Local CPU ONNX vector embeddings with full source code ownership and explainable scoring audits.
OPERATIONAL EXCELLENCE

The Merchandiser Workbench: 8 Production Screens

Built in Django & Wagtail mounted directly inside the FastAPI engine (one process, one port, one unit). Merchandisers never touch terminal configs:

🔒https://anvi-search.internal/workbench/relevance-tuningIN-VPC STAGING / PROD
Site:
Store:
Geo / Climate:
👤 Lead Merchandiser
Relevance Algorithm Tuning: BM25 + Dense Vector RRFRRF k=60 ACTIVETEST PASSEDHybrid Weights & Neural ControlsSparse BM25 Weight (Lexical SKU Precision)0.65Dense Vector Weight (Semantic Conceptual)0.45Neural Cross-Encoder (Reranker)ms-marco-MiniLM-L6-v2 (+0.035 nDCG)query: "waterproof running shoes"Candidate Scoring Verdict (Top 4 of 1,280 Results)RANKSKU / PRODUCT NAMEBM25VECTORRRF SCOREVERDICT#1Arc'teryx Therme Heavy Down Parka📍 NYC Climate (34°F) +300 · 👤 Order Loyalty: Arc'teryx +250 · Size LSKU-ARC-7721 · $850.00 (In Stock Chicago & NYC DC)16.400.9120.0418GEO+RRF#2Speedcross 6 ClimaSalomonSKU-SL-4412 · Salomon12.400.8620.0315FUSED#3Ghost 15 Shield Water-ResistantSKU-BR-7801 · Brooks10.150.8410.0298BOOST#4Hierro V7 All-Terrain VibramSKU-NB-2391 · New Balance8.950.8200.0284ORGANIC
SCREEN S1 · PRODUCTION WORKBENCH

Relevance & Availability Weighting

Tune sparse BM25 vs dense vector weights, and configure In-VPC Contextual Availability & Compatibility Multipliers.

✦ Operational Merchandising Capabilities

  • Select search paradigm: Pure BM25, Pure Vector, or Hybrid Reciprocal Rank Fusion (RRF)
  • Geo-Location & Climate Intent: Query 'jacket' in New York prioritizes heavy down parkas; in Dallas it prioritizes lightweight fleece
  • Customer Order History Affinity: Past brand loyalty (e.g. 3 Arc'teryx orders) and saved size (Size 10.5) dynamically boost matching SKUs
  • Contextual Availability Multiplier: Boost parts compatible with customer's installed-base equipment (+500 pts)
  • B2B Contract Entitlements: Prioritize contract-authorized catalog items and bury unauthorized SKUs
  • One-click opt-in toggle for ms-marco-MiniLM-L6-v2 neural cross-encoder (+0.035 nDCG)

⚙️ Engine Execution Mechanics

Postgres persists weighting profiles. Evaluated during Stage 07 (RRF Fusion) by intersecting customer asset context with compatibility graphs inside your VPC.

🔄 Endeca Migration Parity

Replaces Dgraph relevance modules while matching Coveo-class availability relationship boosting — completely on-premise.

INTERACTIVE ARCHITECTURE DEEP DIVE

End-to-End Retrieval & Ingestion Pipelines

Inspect every microsecond of the 11-stage query execution lifecycle and the three multi-speed catalog ingestion lanes:

11-STAGE QUERY LIFECYCLE

The Deterministic 11-Stage Query Pipeline

Order is the design: each stage shapes what follows it. Click any stage below to inspect its latency budget, data inputs/outputs, and underlying engine technology.

01
🛡️
Request Intake & Context Binding0.2 ms
Authentication, Scope & Header Sanitization
02
✂️
Parse & Dimension Splitting0.4 ms
Unicode Normalization & Alphanumeric Parsing
03
📖
Query-Side Synonyms Expansion0.6 ms
Equivalence & One-Way Merchandiser Synonyms
04
🎯
Landing Page & Trigger Resolution1.1 ms
Multi-Site Inheritance Tree & Depth-Before-Breadth
05
⚖️
Business Merchandising Rules0.9 ms
Dynamic Boost, Bury, Pin & Redirection
06
Concurrent Dual-Branch Retrieval9.7 ms
BM25 + Vector Retrieval with SKU Field Collapse
07
🔄
Application-Layer RRF Fusion1.8 ms
RRF Fusion & Contextual Availability Multipliers
08
🧠
Opt-In Neural Cross-Encoder Reranking82.5 ms
Deep Cross-Attention Relevance Scoring
09
🗂️
Guided Navigation & Facet Dimension Tree3.5 ms
Disjunctive Multi-Select Facet Computation
10
🔍
Speller & Zero-Result Constraints Relaxer0.8 ms
Failsafe Intelligence on Empty Match Sets
11
🚀
Layout Presentation & Diagnostic Egress0.3 ms
Dynamic Swatch Promotion & Store BOPIS Overlay
STAGE 06 OF 11

Concurrent Dual-Branch Retrieval

BM25 + Vector Retrieval with SKU Field Collapse

Latency Budget9.7 ms
Stage Architectural Responsibility:

Executes dual retrieval concurrently: Branch A runs BM25 (Solr edismax) for exact SKU precision; Branch B runs dense vector k-NN (384-dim e5-small). Executes Solr parent-level field collapse: matches child SKU attributes (size, color, spec) while rolling up to 1 clean parent card.

📥 Inputs & Pre-Conditions
  • Rule AST & site_id filter
  • Dense Query Vector
  • Parent-Child SKU Grouping Spec
📤 Outputs & Artifacts
  • Top-100 Lexical Candidates + Scores
  • Top-100 Vector Candidates + Cosine
  • Collapsed Parent Match Groups
Production Technology Stack & Algorithms:
Solr edismax Lexical Corepgvector / Dense HNSW VectorSolr Field Collapse Engine
EVENT-DRIVEN INGESTION ARCHITECTURE

Event-Driven Kafka Streaming vs. Brittle Push/Pull Crawlers

How Anvi Search ingests enterprise catalog mutations: replacing table-locking synchronous crawls with Apache Kafka event streaming, alongside pluggable connectors for legacy systems.

❌ THE LEGACY PUSH/PULL & SYNC BOTTLENECK

Batch crawlers (Endeca CAS / Forge / scheduled JDBC) poll databases every 30–60 min, causing query spikes, database table locks, and stale out-of-stock items. Mid-crawl network errors abort the run, forcing 4-hour restarts.

⚡ THE ANVI KAFKA EVENT-DRIVEN STREAMING ROUTE

Recommended flagship route: ERP/PIM systems publish granular domain events to partitioned Kafka topics. Zero table locks, sub-20ms stock propagation via in-memory overlays, and instant replayability via consumer offset rewind.

Lane 1Sub-20ms

Real-Time Kafka Hot-Path (CDC)

Sub-20ms Continuous Propagation

Lane 2Continuous

Async Embedding & Text Mutations

Continuous Micro-Batch (1-3s Freshness)

Lane 3Stream

Shadow Core Baseline & Stream Replay

Stream Replay / Zero-Downtime Swap

Lane 1: Real-Time Kafka Hot-Path (CDC)

Topic: catalog.inventory-updates & catalog.price-changes
RECOMMENDED STREAMING ROUTE

Captures instant row-level database mutations (stock drops, flash sale discounts, BOPIS store inventory) and streams them via Kafka into an in-memory document overlay with zero search table locks.

1
1. Transactional Mutation in ERP / OMS / POS (Debezium CDC)
2
2. Partitioned Kafka Topic: catalog.inventory-updates (Key: sku_id)
3
3. In-Memory Store Bitset & DocValues Overlay Mutation
4
4. Storefront Real-Time Availability & Price Binding (< 20ms total)
Peak Throughput:25,000 events / sec
Propagation Latency:< 20ms end-to-end
Search Lock Isolation:Zero-Lock In-Memory Overlay
Rollback Contract:Kafka offset rewind
🔌 Pluggable Ingestion Connector Interfaces:Kafka streaming is recommended, with pluggable adapters for existing legacy architectures:
⭐ Kafka Streaming (Recommended)Debezium CDC (Oracle / Postgres / Mongo)REST Bulk Push APIS3 / SFTP Batch Pull Feed
MIGRATION & PARITY BLUEPRINT

Oracle Endeca vs. Anvi Search Architecture

How Anvi Search acts as an empirical migration instrument — replacing legacy MDEX components while preserving 100% of your business rules and storefront contracts.

LEGACY ARCHITECTURE

Oracle Endeca (MDEX 11.3)

01
Dgraph Engine (C++ In-Memory Index)

Single-threaded query processing bound to proprietary MDEX binary. Vulnerable to memory fragmentation.

02
Forge & Dgidx Pipeline

Brittle XML configuration files running multi-hour monolithic batch crawls.

03
Experience Manager (EAC Cartridges)

Proprietary XML rule engine for boost/bury, slotting, and dimension precedence.

04
Navigation State Query Protocol (N, Ne, Ntt, Ntk)

Proprietary query parameters embedded in ATG JSP / React storefront form handlers.

⚡ ZERO REWRITE TRANSLATOR
MODERN IN-VPC REPLACEMENT

Anvi Search Platform

01
Stateless In-VPC Hybrid Core

Concurrent multi-threaded C++ / Rust core executing dual BM25 + Vector HNSW retrieval.

02
3-Lane Ingestion (Batch + Sub-20ms CDC)

Continuous event streaming via Kafka / Debezium with in-memory delta overlays.

03
Dynamic Cartridge Engine + Resolution Inspector

1:1 drop-in XML import with real-time audit tracing and explainable rule collision visualizer.

04
Drop-In MDEX Query Protocol Translator

Transparently accepts legacy Endeca query parameters and maps them to hybrid plans.

MIGRATION PARITY INSPECTOR: COMPONENT 01

Dgraph Engine Stateless In-VPC Hybrid Core

⚠️ Legacy Endeca Bottleneck:

Unpredictable latency spikes under heavy concurrent faceting.

✓ Anvi Search Advantage:

Sub-10ms predictable p99 latency across millions of SKUs with horizontal auto-scaling.

AGENTIC AI & MCP PROTOCOL ARCHITECTURE

Model Context Protocol (MCP) Grounding Pipeline

How autonomous AI shopping agents interact with Anvi Search via standard MCP tool contracts — ensuring 100% factual accuracy and zero telemetry egress.

010.2ms

Dynamic Tool Contract Discovery

AIAnvi
021.1ms

Agent Tool Execution Call

AIAnvi
034.5ms

Hallucination Shield & In-VPC Verification

AnviPrivate
040.3ms

Grounded Citation Payload Delivery

AnviAI
STEP 01 PROTOCOL INSPECTOR

Dynamic Tool Contract Discovery

AI Shopping Agent (Claude / Cursor)Anvi In-VPC MCP Gateway

The AI Agent queries the MCP gateway to discover active catalog dimensions, available tools, and parameter schemas dynamically without hardcoded prompting.

GET /mcp/tools -> Tool RegistryJSON PAYLOAD
{
  "tools": [{
    "name": "catalog_search",
    "description": "Hybrid vector + BM25 catalog search with parametric facet filtering.",
    "parameters": {
      "query": { "type": "string" },
      "filters": { "type": "object" }
    }
  }]
}
NEXT-GENERATION ENTERPRISE ARCHITECTURE

Complete 28-Subsystem Architecture Blueprint

Explore every component across Agentic AI, Model Context Protocol (MCP), Hybrid RAG, 10-Stage Pipeline, 3 Ingestion Lanes, and In-VPC Isolation.

✓ Production Shipping< 1 ms

Model Context Protocol (MCP) Server (GET /mcp/tools)

Standardized tool-calling contract allowing autonomous AI agents (Claude, OpenAI, Gemini) to query catalog facets, search products, and inspect inventory.

Industry-First: Native MCP Server built directly into the search engine core.
Endeca Parity: Not Supported in Endeca (Pre-AI Era Architecture)
✓ Production Shipping< 2 ms

Agent Readiness & Grounding Shield (GET /readiness)

Guarantees zero AI hallucination by evaluating catalog attribute completeness and generating verified citation proofs for LLM outputs.

Hallucination Shield: Mathematically guarantees LLM agent answers are grounded in real SKU facts.
Endeca Parity: Not Supported in Endeca
✓ Production Shipping3.8 ms

Graph-RAG Relational Knowledge Traversal

Traverses product compatibility matrices, variant trees, and bundle graphs to resolve relational questions.

Graph-RAG: Connects parts, accessories, and compatibility graphs inside the search pipeline.
Endeca Parity: Custom ATG/Endeca Relational Cartridge Feeds
✓ Production Shipping1.5 ms

Corrective RAG (CRAG) & Confidence Filtering

Evaluates retrieval confidence across hybrid branches; automatically filters out low-relevance items before feeding generation contexts.

Self-evaluating retrieval that guards against low-quality RAG contexts.
Endeca Parity: Not Supported in Endeca
⚙ Roadmap2.1 ms

In-Session Real-Time Intent Drift Tracking

Dynamically updates user intent vectors within the active shopping session based on click paths without requiring persistent PII.

Real-time vector drift tracking that personalizes without privacy invasion.
Endeca Parity: Endeca Experience Manager Profile Rules (Static only)
⚙ Roadmap18.5 ms

Multimodal Vision-Language Search (SigLIP / CLIP)

Joint vision-language embedding space enabling visual similarity search combined with text attributes.

Unified vision + text semantic space.
Endeca Parity: Not Supported in Endeca
✓ Production Shipping0.4 ms

Query NLU & Token Normalization

Deep query normalization, Unicode case-folding, punctuation normalization, and token-level accent stripping.

Endeca Parity: Endeca Pipeline Tokenizer & Character Mapping
✓ Production Shipping0.6 ms

Query-Side Synonyms Engine

Merchandiser-editable one-way and two-way equivalence synonym expansion running before rules evaluate.

Endeca Parity: Endeca Thesaurus (Equivalence & One-Way Sets)
✓ Production Shipping1.2 ms

NLP Intent & Entity Decomposition

Deconstructs compound queries into structured attribute filters (colors, sizes, brands, price boundaries).

Endeca Parity: Endeca Automatic Category & Dimension Value Recognition
✓ Production Shipping0.8 ms (zero results only)

Vocabulary-Aware Speller & Did-You-Mean

Spelling suggestion that validates against index vocabulary first, firing only when strict matches fail.

Endeca Parity: Endeca Did You Mean (DYM) & Automatic Spell Correction (Aspell)
✓ Production Shipping9.7 ms

Lexical BM25 / Edismax Engine

High-throughput Lucene/Solr inverted index retrieval with configurable multi-match modes.

Endeca Parity: Endeca Dgraph Text Search Engine (MatchMode)
✓ Production Shipping14.2 ms

Dense Vector Semantic Search (kNN)

384-dimensional dense semantic embedding retrieval powered by quantized e5-small-v2 ONNX models.

Endeca Parity: Not Supported in Endeca (Custom Extension)
✓ Production Shipping16.5 ms

Cross-Lingual Retrieval Engine

Retrieves English product catalog records from queries submitted in Spanish, French, or other languages.

Endeca Parity: Not Supported in Endeca (Required Separate Dgraph per language)
✓ Production Shipping1.8 ms

Application-Layer RRF Hybrid Fusion

Reciprocal Rank Fusion in the application tier (Lexical 1.0 / Vector 0.4) preserving sort stacks.

Endeca Parity: Not Supported in Endeca (Endeca only supported Stratified Ranking)
✓ Production Shipping115 ms (Opt-in)

Cross-Encoder Neural Reranker (ms-marco-MiniLM-L6-v2)

Deep sequence-pair cross-attention reranker providing +0.035 nDCG boost.

Endeca Parity: Not Supported in Endeca
⚙ Roadmap8.0 ms

Learning to Rank (LTR / Click-Feedback)

Machine learning ranking (LambdaMART) trained on verified live click, add-to-cart, and purchase streams.

Endeca Parity: Not Supported in Endeca
✓ Production Shipping1.1 ms

Landing Page & Trigger Resolution (Depth > Breadth)

Deterministic page and cartridge resolution walking the navigation state hierarchy.

Endeca Parity: Endeca Experience Manager (Page / Navigation State Triggers)
✓ Production Shipping0.9 ms

Merchandising Rules Engine (Boost, Bury, Pin)

Layered business rules modifying product elevations, score multipliers, redirects, and banner slots.

Endeca Parity: Endeca Business Rules & Boost/Bury Actions
✓ Production ShippingWorkbench UI

The Resolution Inspector (Visual Debugger)

Interactive diagnostic workbench answering Why is this page showing? with complete candidate traces.

Endeca Parity: Not Available in Endeca (Endeca required reviewing XML EAC logs)
✓ Production Shipping3.5 ms

Guided Navigation & Dynamic Faceting

Multi-select facet counts, hierarchical dimension trees, and dynamic facet reordering.

Endeca Parity: Endeca Guided Navigation & Dynamic Dimension Refining
✓ Production Shipping4.2 ms

Negotiable Retrieval on Zero Results

Analyzes binding constraints when queries return 0 results and suggests relaxed alternatives with live counts.

Endeca Parity: Not Supported in Endeca
✓ Production Shipping0.4 ms

Sub-Millisecond Typeahead & Autocomplete

Ultra-fast prefix search honoring tenant, entitlements, and merchandiser exclusions.

Endeca Parity: Endeca Dimension Value Search & Typeahead
✓ Production ShippingPipeline Sync

Lane A: Batch Catalog Sync & Blue-Green Reindex

12-step full catalog reindex with validated configsets, monotonic guards, and node parity validation.

Endeca Parity: Endeca Baseline Index Pipeline (Forge / Dgidx / EAC Baseline Update)
✓ Production Shipping< 20 ms CDC

Lane B: Real-Time Price & Inventory Streaming (CDC)

Microsecond price, stock, and status updates via high-watermark PostgreSQL streaming.

Endeca Parity: Endeca Partial Updates (Delta Feed / EAC Partial Index Update)
✓ Production ShippingAsync Graph Feed

Lane C: Dynamic Graph & Product Relations

Calculates and updates product bundles, cross-sells, and compatible accessories.

Endeca Parity: Custom ATG/Endeca Relational Cartridge Feeds
✓ Production ShippingRealtime Logs

Query Health & Fallback Analytics

Monitors search fallback rates, zero-result frequency, and shallow rerank pool warnings.

Endeca Parity: Endeca Search Analytics & Log Server
✓ Production ShippingAtomic TX

Atomic Append-Only Audit Trail

Complete, immutable record of every merchandising change, rule edit, and API key action.

Endeca Parity: Not Supported in Endeca (Endeca overwritten XML without version history)
✓ Production ShippingZero Egress

100% In-VPC Isolation & Data Protection

Runs entirely within your cloud VPC or bare-metal environment with zero external telemetry egress.

Endeca Parity: Endeca On-Premise Deployment Model
Production ShippingLatency Budget: < 1 ms

Model Context Protocol (MCP) Server (GET /mcp/tools)

Standardized tool-calling contract allowing autonomous AI agents (Claude, OpenAI, Gemini) to query catalog facets, search products, and inspect inventory.

Innovation Breakthrough: Industry-First: Native MCP Server built directly into the search engine core.

Technical Execution & Design Guarantees

  • Dynamically generates tool parameters and JSON Schema directly from the active catalog attribute registry.
  • Strict tenant and credential encapsulation preventing agent token leakage.
  • Allows multi-agent shopping assistants to perform complex facet filtering without hardcoded assumptions.
Oracle Endeca Drop-in Equivalent
Not Supported in Endeca (Pre-AI Era Architecture)
Enables drop-in migration without retraining business merchandisers or restructuring product catalog feeds.
RIGOROUS EMPIRICAL BENCHMARKS

Measured on 4,998 Real Amazon Products (5,482 Human ESCI Judgements)

Search quality must be scientifically measured against real catalog data, not assumed. Here are the empirical measurements taken on running Anvi Search nodes:

Retrieval ConfigurationnDCG@10 ScoreLatency (Laptop)Production Status & Architectural Role
Lexical Alone (BM25 edismax)0.5273.5 msFast exact token, SKU, and keyword matching.
Vector Alone (e5-small k-NN)0.4924.8 msSemantic recall; loses on exact SKUs when unassisted.
Anvi Hybrid Fusion (RRF)0.533 – 0.5379.3 msProduction Default: Zero-loss semantic + exact SKU recall.
Hybrid Fusion + Cross-Encoder Rerank0.569 (+0.035 nDCG)82.5 msOpt-In: Maximum precision reranking for top-25 candidates.
OPEN ARCHITECTURE & KNOWLEDGE SHARING

Passionate About Search Architecture & Engineering?

Anvi Search is an open hobby and learning research endeavor. I love connecting with fellow software engineers, search architects, and technology enthusiasts to exchange ideas, discuss retrieval patterns, and share learnings.

Connect & Say Hi →Browse 108 Technical Guides