Vector Search Without Ops

Managed vector search and retrieval for engineering teams. 24 ms p99 at a billion vectors, recall you can measure, and no index for your team to babysit.

[ The Problem ]

Why Your Own Index Hurts

Three things go wrong with a self-run vector index, and they all go wrong at the same size.

Latency You Cannot Explain

p99 triples the week you cross fifty million vectors, and nobody on the team owns the index.

Recall Nobody Measures

Approximate search quietly drops the right answer. You hear about it from a user, not a dashboard.

A Database You Never Hired For

Shards, rebuilds, replica lag, memory ceilings. Two engineers end up running a store nobody set out to build.

[ The Solution ]

Everything the index needs

One retrieval layer with a published latency number, a recall report and nothing for you to operate.

Queries / Second
1,240 +70%
SUNMONTUEWEDTHUFRISAT
Index Growth · Vectors Written
Jan 2026 +47.2K
Feb 2026 +38.9K
Mar 2026 +52.4K

Managed HNSW and DiskANN Indexes

We pick the index family, shard count and replica layout, then keep them inside your latency budget.

Hybrid Queries
7d 30d 1m 6m
124 K queries / 7d
sunmontuewedthufrisat

Hybrid Retrieval in One Query

Dense vectors, keyword BM25 and metadata filters resolved together, in a single round trip.

Audit Logging 256-bit Encryption Zero Trust Regional Isolation Key Rotation Audit Logging 256-bit Encryption Zero Trust Regional Isolation Key Rotation
End-to-End Encrypted Audit-ready Data Residency mTLS Identity PII Redaction End-to-End Encrypted Audit-ready Data Residency mTLS Identity PII Redaction

Private by Default

Single-tenant indexes, encryption at rest and in transit, and a region you choose.

Embedding Backlog 48 Shards
28% Avg Recall
85% Optimal 84% Partial

Embedding Sync Built In

Point us at Postgres, S3 or a queue. New rows are embedded, indexed and searchable in seconds.

[ Our impacts ]

Numbers we publish and are held to

24ms

p99 Query Latency

99%

Recall @10, Scale Tier

12B+

Vectors Indexed

40K/s

Peak Queries Served

[ how it works ]

How it works

From zero to production in three steps

Step 1

01 - Load

Push vectors over REST or gRPC, or point us at Postgres, S3 or Kafka. We build the index.

  • Bulk import at up to two million vectors a minute

  • Schema and dimension checks before anything is indexed

  • First query answered inside five minutes

Step 2

02 - Query

One endpoint for approximate search, metadata filters and keyword hybrid, with clients for Python, TypeScript and Go.

  • Per-query latency, recall and cost in the response headers

  • Filters applied inside the index, not after it

  • Read replicas in the regions your users are in

Queries / Second
1,240 +70%
SUNMONTUEWEDTHUFRISAT

Step 3

03 - Tune

Give us a recall target and a latency budget. We tune the index against your own traffic and show the working.

  • Weekly recall report against exact ground truth

  • Automatic rebuilds with no read downtime

  • A named engineer on the Scale tier and above

Recall Target 0.97
Measured 0.993

[ Testimonials ]

What engineering teams say

Trusted by teams shipping AI products at scale.

"We moved 300 million embeddings onto Nodeform over a weekend and our p99 dropped from 180 ms to 22 ms. Nobody has touched an index since."

Nadia Farouk

Nadia Farouk

Staff Engineer, Arclight Data

"The recall report is the part I did not know I needed. We can finally tell product why an answer was missed."

Callum Reid

Callum Reid

VP Engineering, Cobalt Works

"We went from 40 million vectors to 1.4 billion in nine months. No rebuild, no migration weekend, no config change."

Wei Zhang

Wei Zhang

ML Platform Lead, Tessell Labs

"Hybrid filters in the same query killed a whole service for us. One endpoint now does what three of ours used to do badly."

Joana Ribeiro

Joana Ribeiro

Head of Search, Kestrel AI

[ blogs ]

Latest from our blog

Benchmarks, postmortems and notes on running vector search in production.

View all posts