The Postgres Search Wars: How ParadeDB Closed the Gap Against PlanetScale’s TIN in Record Time

By Ming Ying
Published: October 1, 2026
Main Facts: The Battle for Postgres Full-Text Search Supremacy
The landscape of relational database search capabilities shifted dramatically following a high-stakes performance rivalry between two prominent players in the Postgres ecosystem. Just weeks after cloud database giant PlanetScale unveiled TIN—a proprietary, high-performance full-text search extension for Postgres—ParadeDB responded with a masterclass in aggressive engineering.
PlanetScale’s initial launch benchmarking claimed eye-popping performance wins, outperforming ParadeDB (version 0.25) by at least an 8x margin across a suite of tests focused on BM25-ranked text search and document counts. PlanetScale attributed this blazing speed to a foundational architectural pivot: leveraging Postgres’s internal ctid physical storage pointers directly as document identifiers, thereby sidestepping the mapping overhead traditional search libraries require.
Rather than dismissing the benchmarks or engaging in defensive posturing, the ParadeDB engineering team put on their optimization hats. Within two weeks, using the exact same hardware, datasets (the 150-million-document StackExchange benchmark), and harness configurations, ParadeDB eliminated the performance deficit. Remarkably, they achieved this parity without altering their underlying document identification architecture—proving that architectural divergence was not the sole determinant of search throughput.
Chronology: Two Weeks of Intense Optimization
The timeline of this engineering sprint highlights the rapid iteration cycles typical of modern database infrastructure development:
- Two Weeks Ago: PlanetScale publicly launches TIN, releasing benchmark data that highlights overwhelming performance dominance over ParadeDB 0.25 on BM25 Top-K queries and aggregate
COUNToperations. They credit their success to bypassing ID translation layers through the use of Postgresctidfields. - Days 1–5 (The Investigation): The ParadeDB team adopts the official
paradedb/benchmarkersuite to run localized diagnostics on a smaller 28.7-million-document Hacker News dataset. They discover that random-access bottlenecks—specifically regarding "fieldnorms" (document length metadata)—are severely degrading query efficiency. - Days 6–10 (Algorithmic Shifts): Profiling reveals that disjunction queries with numerous terms suffer from excessive overhead under the standard Blockmax WAND loop. ParadeDB implements a hybrid MAXSCORE pruning path with a dynamic selection heuristic.
- Days 11–14 (Syntax and Configuration Audit): The team scrutinizes the original PlanetScale benchmark parameters. They identify and correct an unqualified syntax discrepancy in ParadeDB’s query parser and uncover subtle trade-offs concerning TIN’s "dense-term elision" scoring shortcuts.
- Today: ParadeDB cuts release candidate
0.26.0-rc.2, featuring major architectural optimizations, and publishes a transparent, dual-configuration benchmark breakdown for the community.
Supporting Data: Dissecting the Performance Levers
To understand how ParadeDB closed an order-of-magnitude gap in just fourteen days, one must examine the specific engineering interventions deployed across the codebase.
1. Eliminating Random Access via Per-Term Fieldnorms
In BM25 scoring, engines normalize scores based on document lengths, encoded as tiny values called "fieldnorms." Traditionally, Tantivy (the search library underpinning ParadeDB) stored fieldnorms globally in an array indexed by internal document IDs (DocId). While memory-mapped storage makes this efficient in isolation, running it within Postgres block storage triggered roughly 1,500 distinct page accesses per query.
ParadeDB refactored this layout by storing a dedicated fieldnorm array alongside each individual postings list, ordered identically to the postings’ DocId values.
- The Result: Sequential reading eliminated scattered memory lookups, causing fieldnorm page accesses to plunge from 1,500 down to just 30.
- The Tradeoff: A modest 9% increase in index storage footprint on the Hacker News corpus, as document fieldnorms are duplicated across distinct terms.
2. Upgrading to Blockmax MAXSCORE Pruning
While single-term queries saw immediate relief, complex multi-term disjunctions (e.g., termA OR termB OR termC...) remained bound by algorithmic overhead. Profiling demonstrated that the standard Blockmax WAND loop was spending excessive CPU cycles deciding what chunks of postings to skip.
Borrowing concepts from modern search architecture evolutions (such as Apache Lucene’s dynamic pruning strategy), ParadeDB introduced a MAXSCORE execution path. The engine now dynamically evaluates query shapes:

- MAXSCORE is deployed for disjunction queries containing three or more terms with dense postings, minimizing overhead.
- WAND remains active for all other query shapes.
- The Result: On a 10-term disjunction query over the Hacker News dataset, p50 latency dropped by roughly 6x, p95 latency plummeted by 8x, and ParadeDB surpassed TIN’s throughput twofold.
3. Untangling Benchmark Anomalies: Syntax and Elision
ParadeDB’s audit of PlanetScale’s benchmarks uncovered two critical discrepancies that inadvertently skewed the initial results:
- The Query Parser Oversight: TIN benchmarks evaluated ParadeDB using its mini-query language (
@@@) without explicit field qualification (e.g., searching<query>instead of<field>:<query>). This forced ParadeDB to scan multiple indexed text columns simultaneously, whereas TIN targeted a single column. Shifting tests to native qualified operators (|||for disjunctions) leveled the playing field. - Dense-Term Elision: TIN achieves high performance on common terms (like "the" or "is") by implementing "dense-term elision"—skipping scoring entirely for terms appearing in over 10% of the corpus. While this dramatically accelerates execution, it sacrifices exact BM25 compliance. When applied to synthetic corpora generated by sampling consecutive word spans (such as "is it" or "to a"), elision caused TIN to bypass scoring work that ParadeDB computed exactly. When tested against exact BM25 matching configurations (
dense_ratio=2), the playing field normalized significantly.
Official Responses and Industry Implications
The database community has closely watched the exchange, viewing it as a healthy exercise in open competition.
PlanetScale’s initial framing positioned its proprietary ctid-centric architecture as a silver bullet for full-text search inside relational systems. By utilizing physical tuple pointers, TIN avoided translation layers between search identifiers and relational blocks.
However, ParadeDB’s counter-analysis challenges the narrative that document identifier choice is an absolute architectural prerequisite for speed. ParadeDB’s maintainers emphasize that dense, sorted, unique u32 integer IDs (DocId) offer unmatched compression and form a vital bridge to columnar storage.
"TIN is great at BM25 scoring and document counting, but that’s just the tip of what makes up a search engine like Elasticsearch," the ParadeDB team noted. "For the ‘rest of search,’ you need a columnar representation. If TIN decides to make one, we suspect they’ll have to pay the same
ctid/DocIdtranslation cost (but in the reverse direction)."
Furthermore, the decision to build upon Tantivy rather than writing a search engine from scratch has been fully vindicated in the eyes of ParadeDB’s contributors. Tantivy brings a decade of battle-tested development, rich extensibility, and rapid adaptability—qualities that enabled a two-week turnaround on complex scoring and pruning optimizations.
Conclusion: The Road Ahead for Postgres Search
The rivalry between PlanetScale’s TIN and ParadeDB underscores a booming market demand: developers increasingly want powerful, native full-text search directly embedded within their relational databases.
For ParadeDB users, the fruit of this high-pressure optimization sprint arrives next week with the stable release of version 0.26.0. While a reindex will be required for existing tables to inherit the full suite of layout improvements, the upgrades are entirely backward-compatible and promise substantial latency reductions across production workloads.
As both closed-source and open-source platforms continue to push the boundaries of what Postgres can achieve, end-users are the ultimate beneficiaries. With Part II of ParadeDB’s performance series promising a deep dive into aggregate COUNT optimizations, the Postgres search wars are clearly only just beginning.
