September 29, 2026

Breaking the Seven-Year Silence: How PostgreSQL 19 Finally Fixed Its Simplified Chinese Localization

breaking-the-seven-year-silence-how-postgresql-19-finally-fixed-its-simplified-chinese-localization

breaking-the-seven-year-silence-how-postgresql-19-finally-fixed-its-simplified-chinese-localization

For generations of database administrators (DBAs) operating PostgreSQL environments across China, an unwritten, ironclad rule governed every fresh installation: when configuring the database, bypass the system prompt, skip the zh_CN locale entirely, and default straight to en_US.

This directive was never mere superstition. Rather, it functioned as cumulative scar tissue borne by an entire generation of engineers who learned the hard way that native language support within the database ecosystem was fundamentally broken. To receive coherent diagnostic messages from a server operating on Chinese soil, users were effectively forced to disable their own language.

As of the upcoming release of PostgreSQL 19, that era has officially ended.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

Database architect Ruohang Feng has successfully overhauled PostgreSQL’s Simplified Chinese localization from the ground up. Spanning six major software versions—from PostgreSQL 14 through 19—the monumental effort encompassed 28 distinct message catalogs, 67,487 unique strings, and achieved 100% translation coverage. Veteran PostgreSQL contributor Peter Eisentraut formally merged the new catalogs into the upstream translation repository, ensuring they will ship natively with PostgreSQL 19.


Main Facts: The Anatomy of a Long-Standing Crisis

Open-source relational database management systems rely heavily on National Language Support (NLS) machinery, powered by official gettext message catalogs, to translate server-emitted English strings into localized equivalents. However, the mere presence of the NLS framework does not guarantee operational utility.

PostgreSQL enforces a strict packaging threshold: no language is permitted to ship with an official release unless its message catalogs clear an 80% completion rate. Anything falling below that baseline is automatically excluded from distribution.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

By mid-2024, Simplified Chinese localization sat perilously at 66% completion, while Traditional Chinese lagged at 70%. Consequently, both variations faced imminent deletion from the PostgreSQL 19 release cycle.

The vertical breakdown of the catalogs painted an even grimmer picture. Out of 28 core Simplified Chinese catalogs, only 9 cleared the 80% packaging line. Most critically, the primary postgres server catalog—which houses over 6,800 strings and serves as the genetic origin for nearly every operational ERROR encountered by developers—hovered at a dismal 61% completion rate. The foundational libpq client library catalog languished at a mere 12%.

When gettext cannot locate a localized string, it defaults to the English source text. In practice, this meant users were greeted with fractured error logs where the first half of a diagnostic string appeared in Chinese while the latter half defaulted to English. Different administrative tools utilized entirely distinct terminologies for the same underlying database concepts, making severe production incidents exceedingly difficult to diagnose.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

Chronology: Seven Years of Stagnation Reversed in Seven Days

An inspection of the file headers within the upstream codebase reveals the timeline of neglect. The primary Simplified Chinese catalog file (src/backend/po/zh_CN.po) remained completely frozen for seven years and three months, spanning eight major releases from PostgreSQL 12 through 19. While minor, sporadic patches touched isolated client components like psql or pg_ctl, the massive server-side core remained entirely untouched. Approximately 2,660 new strings introduced via features like logical replication, Just-In-Time (JIT) compilation, parallel query execution, and asynchronous I/O possessed entirely blank Chinese equivalents. Furthermore, roughly 1,000 legacy translations no longer matched their corresponding source IDs (msgid) due to upstream code refactoring.

The turnaround began in earnest in September 2026:

  • September 11: Ruohang Feng formally emailed the pgsql-translators mailing list, volunteering to assume complete stewardship over the Simplified Chinese catalogs.
  • September 13: The first comprehensive batch covering all 28 catalogs for PostgreSQL 19—totaling 12,702 strings—was submitted alongside a tracking issue on Redmine.
  • September 17: The scope expanded aggressively. Every catalog across six major branches (PostgreSQL 14 through 19) was updated, resulting in 162 modified files comprising 67,487 strings at 100% translation coverage. All files passed strict msgfmt --check --check-format linting tests.
  • September 18: Peter Eisentraut formally merged the submissions into the official PostgreSQL 19 translation repository.

What seven years of community passivity failed to accomplish, a concerted, multi-day engineering effort achieved in precisely one week.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

Supporting Data: The Museum of Mistranslations

Beyond mere absence, the historical catalogs were littered with egregious mistranslations that actively misled administrators during critical operational crises.

Among the most notorious was the translation of the ubiquitous diagnostic prompt Report bugs to, which was rendered literally as "report the bedbugs to" (报告臭虫与). This mistranslation—originating from an initial 2001 submission—survived unmodified across fourteen separate catalogs for a quarter-century.

Another hazardous anomaly involved the translation of out of memory (OOM). Across the legacy codebase, four distinct Chinese renderings competed for dominance. While three were acceptable, the fourth rendered the error as "memory overflow" (内存溢出). In systems engineering, these concepts are fundamentally distinct: an OOM error signifies a clean operational failure where the process requests memory and is denied by the kernel, whereas a buffer overflow represents a critical memory corruption vulnerability where data writes past allocation boundaries. During a high-stress production outage, confusing these two failure states sends engineers chasing entirely wrong hypotheses—either tuning work_mem parameters or debugging source code.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

Other errors compromised database security and diagnostics directly. In the pg_ctl utility, the error string invalid binary "%s": %m entirely stripped the %m format specifier in the Chinese translation, depriving administrators of critical operating system error numbers (errno). Meanwhile, an ecpg catalog error inadvertently swapped format specifiers from %s to %1$s and %2$s, causing the localized string to attempt reading non-existent secondary arguments at runtime.


Official Responses and Methodology

The resurrection of the Simplified Chinese catalogs was not executed via raw, unvetted machine translation. Feng leveraged advanced large language models (including variants of Fable and Codex) to accelerate drafting, but implemented a rigorous human-in-the-loop review architecture.

A dedicated web workbench (pgsql.cc/nls) was engineered to facilitate side-by-side human validation of source strings against target outputs. Crucially, translation quality was anchored by an authoritative, domain-specific professional glossary synthesized from Feng’s previous works, which included translating the entirety of the PostgreSQL documentation set (spanning versions 9.0 through 20) and compiling comprehensive PostgreSQL knowledge graphs.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

By freezing terminology rules, modal verbs, and stylistic guidelines prior to mass generation, the project maintained absolute cross-component and cross-version consistency—a feat that previous piecemeal contributors could not sustain.

Upstream maintainers welcomed the overhaul, swiftly clearing the technical barriers that previously threatened to purge Chinese localization from the PostgreSQL ecosystem entirely.


Implications: Open Source Stewardship and Downstream Impact

The successful revival of PostgreSQL’s Simplified Chinese localization carries profound implications for the global database community, open-source governance, and commercial vendors operating within the Chinese market.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

1. The Paradox of Commercial "Domestic" Databases

An investigation into various commercial database products built upon PostgreSQL revealed a sobering reality: numerous local enterprise vendors marketed proprietary databases emphasizing native Chinese support while shipping binaries containing upstream’s unedited, bug-ridden zh_CN.po files down to the exact byte. Despite commanding substantial market shares and leveraging state-backed subsidies, commercial entities failed to contribute back to the foundational open-source project, leaving upstream maintenance to an independent, single-person operation.

2. Preventing Downstream Contamination

Mistranslated technical strings do not remain isolated within localized .po files. They propagate virally into third-party PostgreSQL forks, community blog posts, Q&A forums, technical support tickets, and—increasingly—the training corpora of artificial intelligence models. Left unchecked, corrupted technical definitions risk becoming accepted industry jargon. Rectifying these errors at the upstream source halts the pollution of the broader technical ecosystem.

3. A Call to Open Source Action

Ultimately, the seven-year vacancy of the Simplified Chinese catalog highlights a defining truth of open-source software: features and localization do not materialize via corporate mandate; they require active human stewardship.

PostgreSQL's Chinese Error Messages Were Seven Years Stale. Not Anymore.

As PostgreSQL 19 prepares for general availability, administrators operating in Chinese-speaking environments can finally delete their legacy configuration workarounds. By resetting system locales to zh_CN, users will be greeted not by fragmented syntax or literal references to bedbugs, but by precise, professional database diagnostics.