Google crawled 270 of my articles and indexed none — the body it received was 0 characters

In the ecosystem of modern web development, the Single-Page Application (SPA) has reigned supreme for over a decade. Powered by frameworks like React, Vue, and Angular, SPAs offer lightning-fast client-side transitions and fluid user experiences. Yet, beneath their slick interfaces lies a foundational architectural flaw that continues to blindside developers: search engine crawlers and AI agents do not experience the web the way humans do.
For one solo developer and indie publisher, this architectural blind spot resulted in a shocking discovery on August 13, 2026. After opening Google Search Console and analyzing 427 indexed rows of data, the reality set in: 270 articles had been meticulously written, published, and completely ignored by search engines. Google hadn’t rejected the content because of poor keyword optimization, lack of depth, or weak backlinks. It had rejected them because, from the perspective of a crawler fetching the raw URL, the body of every single article contained precisely zero characters.
This deep-dive technical autopsy explores the anatomy of that failure, the step-by-step diagnostic chronology, the hard data behind the fixes, and the profound implications this holds for the future of JavaScript-heavy publishing platforms.
Main Facts: The Invisible Web of Client-Side Rendering
The core issue facing the developer’s platform—a content-heavy niche site—was a classic symptom of unoptimized Single-Page Application architecture. When a user navigates to an SPA, the server typically responds with a lightweight HTML shell. This shell contains a root DOM node—such as <div id="app"></div>—alongside script tags that point to heavy JavaScript bundles.
Once the browser downloads, parses, and executes those scripts, the JavaScript injects the actual content into the DOM, making the page visible to human visitors.
However, search engine crawlers and AI-search agents (like Googlebot and GPTBot) operate under strict time and resource constraints. While search engine giants maintain that they can execute JavaScript, "will execute" and "will always wait for an application to finish rendering before scoring a page" are two entirely different promises.
When Googlebot fetched the site’s URLs, it received the initial server-side shell: a stark, empty container. Finding no indexable text, metadata, or semantic structure in the raw payload, Google’s algorithms made a logical conclusion: Crawled – currently not indexed.
This status represents a developer’s worst nightmare. It is far more detrimental than a "Discovered – currently not indexed" status, which simply means the crawler hasn’t arrived yet. "Crawled and declined" means the bot actively read the page, evaluated the payload, and judged that it contained nothing worth indexing. Compounding the issue, legacy routing behavior meant that dead URLs were returning a 200 OK status code alongside the homepage shell, resulting in 157 "Soft 404" errors that actively poisoned the site’s search health.
Chronology: The Investigation and Diagnostic Timeline
The path from total obscurity to search engine compliance was not paved with traditional SEO tactics. Instead, it required a rigorous engineering audit divided into distinct phases.
Phase 1: Moving Beyond SEO Myths
When the initial Search Console report revealed 270 unindexed articles, 157 soft 404s, and 95 canonical duplicates, the developer’s first instinct pointed toward conventional fixes: keyword density, article depth, and backlink acquisition. Recognizing that pursuing these paths would waste months on content nobody could discover, the developer pivoted to a golden rule of systems debugging: Stop guessing what the crawler sees, and become the crawler.
Phase 2: The Curl Reality Check
Using a simple command-line HTTP client (curl), the developer requested an article URL utilizing Googlebot’s exact User-Agent string. Saving the raw HTML and checking its length revealed an unvarnished truth: zero characters inside the application body. Repeating the test with OpenAI’s GPTBot yielded identical results. The site was completely invisible not just to traditional search engines, but to the entire emerging ecosystem of AI-driven search platforms. The most humbling realization? That raw HTML response was just one command away the entire time. The developer had successfully published 270 articles before fetching the raw payload even once.
Phase 3: Implementing Server-Side Body Injection
The site already utilized a Netlify edge function for Open Graph (OG) tag injection, meaning it fetched article data from the database to generate social media previews. Leveraging this existing architecture, the developer bypassed client-side rendering bottlenecks for bots.
For human visitors, the site experience remained untouched. For crawlers and bots hitting the edge, the function injected the complete article body, titles, and metadata directly into the initial HTML payload.
To prevent Cross-Site Scripting (XSS) vulnerabilities without a DOM-dependent sanitizer like DOMPurify (which fails in edge runtimes like Deno), the developer built an escape-first renderer directly on the edge. Every character is escaped first, and structural markdown is applied second, rendering XSS structurally impossible.
Phase 4: Eradicating Soft 404s and Fail-Open Hazards
To fix the 157 soft 404 errors caused by legacy WordPress URLs returning 200 OK via wildcard SPA routing rules (/* /index.html 200), the developer implemented strict path validation. URLs that did not match valid database patterns or UUID structures were explicitly routed to throw true 404 Not Found headers.
Crucially, this introduced a high-stakes engineering challenge: fail-open resilience. If the database suffered a three-second latency spike, the application could incorrectly treat "query failed" as "record does not exist," serving false 404s that Google would promptly catalog. To mitigate this, the data helper function was rewritten to explicitly distinguish between an operational query failure and a verified zero-result lookup, safeguarding the site against transient backend hiccups.
Phase 5: Trimming Performance Bloat
With indexing resolved, the developer targeted site performance. A local Lighthouse audit exposed a staggering bottleneck: a single Google Fonts stylesheet (Noto Sans TC) was imposing a render-blocking delay of 2,222 milliseconds, dragging along 98 KiB of data—92 KiB of which consisted of unused CJK Unicode-range subset declarations. This was fixed by dynamically swapping the stylesheet link rel attribute upon load and using system fallback fonts.
Concurrently, an errant manual bundling rule was found to be preloading a heavy 531K PDF vendor bundle on every single page load—including the 404 and newsletter pages. By explicitly returning undefined for those packages and restoring Rollup’s native code-splitting capabilities, the core vendor bundle dropped from 499K to 419K.
Supporting Data: Before and After the Fixes
The scope and efficiency of the remediation efforts are best illustrated through quantitative metrics gathered across the development lifecycle.
The Search Console Breakdown
- Crawled – currently not indexed: 270 rows
- Soft 404s: 157 rows
- Duplicate, Google chose a different canonical: 95 rows
- Total analyzed initial rows: 427
Performance and Bundle Metrics
| Metric | Before Optimization | After Optimization |
|---|---|---|
| Google Fonts Delay | 2,222 ms | < 100 ms (Non-blocking) |
| Module Preloads | 6 (including 531K PDF bundle) | 5 (PDF bundle excluded) |
| Core Vendor Bundle | 499 KiB | 419 KiB |
| Precache Size | 2,398 KiB | 2,319 KiB |
| First-Screen Static Closure | N/A | 6 chunks / 1,461 KiB |
To ensure absolute reliability, every phase of the refactoring was verified using rigorous mutation testing. By deliberately introducing broken code variants (such as removing fail-open protections or altering sanitization logic), the developer confirmed whether the test suites caught the regressions. In multiple instances, automated mutation testing exposed subtle flaws—such as test assertions matching substrings from unrelated parts of the codebase—preventing catastrophic production bugs before deployment.
Official Responses and Industry Implications
While this case study reflects the solo journey of an independent developer, its implications resonate deeply across the broader software engineering and SEO industries.
Web platform engineers and search engine optimization specialists have long debated the true efficacy of dynamic rendering and client-side JavaScript indexing. Major search engine representatives have repeatedly assured developers that modern crawlers are equipped with headless browsers capable of executing JavaScript.
However, real-world case studies like this one expose a dangerous gap between theoretical crawler capabilities and practical execution realities. Search engine bots operate under strict rendering budgets, timeouts, and resource-allocation limits. When an SPA forces a crawler to wait for heavy JavaScript execution frameworks, asynchronous API calls, and client-side hydration just to discover the basic textual body of an article, crawlers frequently time out or deprioritize the payload, resulting in the dreaded "Crawled – currently not indexed" status.
Furthermore, the rise of Retrieval-Augmented Generation (RAG) systems and AI-search agents (GPTBot, ClaudeBot, Perplexity) places an even higher premium on clean, server-side rendered, semantic HTML. AI agents rarely execute client-side JavaScript applications; they consume raw, flat HTML feeds. If an SPA hides its content behind a JavaScript shell, it effectively renders itself invisible to the next generation of AI-driven discovery platforms.
Conclusion: The Modern SPA Content Checklist
Building a content-heavy publication on top of a Single-Page Application framework introduces unique friction that standard development boilerplates rarely address. For engineering teams and indie publishers operating JS-driven sites, avoiding the "ghost site" trap requires adopting a proactive audit protocol:
- Fetch as the Bot: Never trust local assumptions. Run a command-line
curlrequest using a bot User-Agent against your production URLs to inspect the raw HTML payload. If the body is empty, search engines and AI agents see nothing. - Implement Server-Side Edge Injection: Ensure critical article bodies, titles, and metadata are baked directly into the initial server response or handled via edge functions for bot requests.
- Audit Wildcard SPA Routing: Ensure your fallback routing (like
/* -> /index.html) does not return a200 OKstatus code for dead, legacy, or mistyped URLs. Enforce true404 Not Foundresponses to prevent soft 404 indexing penalties. - Guard Against Fail-Open Hazards: When writing database verification layers for custom routing and error handling, ensure transient backend connection hiccups do not masquerade as missing content.
- Embrace Mutation Testing: Do not rely solely on green test suites. Actively break your code to verify that your unit and integration tests are genuinely validating production behavior rather than relying on unrepresentative test fixtures.
Ultimately, the lesson learned from 270 invisible articles is stark: writing exceptional content is only half the battle. In an era dominated by automated crawlers and algorithmic curation, making sure your architecture allows the machine to read what you wrote is the true prerequisite for survival on the modern web.
