Performance Engineering15 min read

Edge Caching for Core Web Vitals: A TTFB-to-LCP Blueprint

A measurement-first guide to deciding whether static delivery, HTML edge caching, edge compute, regional compute, or origin repair can improve the TTFB-to-LCP path.

Too technical? Pick your depth.

Same topic, explained for where you are — from a first-timer to a working specialist.

Edge caching can improve real-user LCP when slow HTML delivery or distant, cacheable critical resources are a material part of the LCP path. It cannot repair late resource discovery, heavy client-side rendering, slow interactions, or layout shifts. Prove the change by comparing equivalent hit, miss, and bypass cohorts by region and percentile, then validate field LCP, not a single warm-cache test.

That distinction matters because “move it to the edge” is an architecture choice, not a performance result. A shared cache can remove repeated origin work. Edge compute can move request logic closer to a visitor. Neither guarantees that the browser discovers the LCP resource promptly, renders it without delay, or remains responsive after the page appears.

This guide provides an Edge Performance Proof Protocol: a route-safety matrix, a request-path latency budget, and a reproducible release test. It is a method for generating project evidence, not a benchmark or a promise of sub-second delivery.

Start with the request path, not the architecture label

An incoming navigation can cross several systems before the browser receives the first byte of HTML:

TEXT
Visitor
  → redirects and connection setup
  → CDN or edge routing
  → shared-cache lookup
  → edge or regional application compute
  → origin, APIs, database, or CMS
  → HTML response
  → critical resource discovery and rendering

Different interventions remove different waits. A CDN can terminate the connection near the visitor yet still forward every HTML request to a distant origin. A cache hit can avoid application and database work, while a miss still pays for the full path. Edge compute can reduce application distance but become slower if it makes sequential calls to data stored in one remote region.

Use these terms precisely:

LayerWhat it changesWhat it does not establish
CDN termination and routingConnection handling and the route toward the applicationThat HTML is cached or application work is nearby
Static asset cacheReuse of images, CSS, JavaScript, and fontsFast initial HTML or correct resource priority
Shared HTML/response cacheReuse of an entire eligible response across visitorsSafety for personalized content or fresh content after an invalidation error
Edge computeRequest logic in a distributed runtimeNearby data, low cold-start cost, or a short backend waterfall
Regional application computeApplication work near its data and integrationsProximity to every visitor
Origin repairFaster database queries, API calls, rendering, and dependency orchestrationShort network distance for a global audience

The right design can combine layers. A marketing page may be statically generated and cached globally; an authenticated dashboard may run in the region that owns its data; its versioned assets may still use a global CDN. Architecture should follow the route's data and correctness requirements, not one site-wide slogan.

Where does TTFB sit inside LCP?

Time to First Byte measures the interval from the start of navigation until the browser receives the first response byte. It includes more than server execution: redirects, connection setup, network travel, CDN behavior, and response delay can all contribute.

TTFB is not a Core Web Vital. The current web.dev TTFB guidance describes 0.8 seconds or less as a rough guide for many sites, while explicitly noting that TTFB is not a Core Web Vital and should be judged by whether it impedes user-facing metrics.

Largest Contentful Paint can be separated into four contiguous parts:

TEXT
LCP = TTFB
    + resource load delay
    + resource load duration
    + element render delay

The official LCP optimization guide recommends diagnosing those parts individually. Lowering TTFB can reduce LCP when the other parts remain controlled. But time saved at the server can disappear into a late image request, render-blocking CSS, a font dependency, a hydration gate, or main-thread work.

Use a second equation to account for the request path:

TEXT
TTFB = redirect
     + connection
     + edge routing
     + cache lookup
     + edge or origin compute
     + backend dependencies
     + response transfer

These are diagnostic buckets, not a universal timing model. Populate them with evidence from the selected platform and actual traffic. Do not insert industry averages and present them as a forecast for your routes.

Google's current Core Web Vitals documentation defines the three field metrics as LCP, INP, and CLS. The recommended “good” thresholds are LCP at or below 2.5 seconds, INP at or below 200 milliseconds, and CLS at or below 0.1. Evaluate them at the 75th percentile, with mobile and desktop segmented, as described in the Web Vitals guidance.

Which delivery pattern fits the measured bottleneck?

Choose the smallest change that can remove the attributed delay. The following matrix is AppWebSeo's editorial decision framework, not a vendor benchmark.

Evidence and route requirementFirst pattern to evaluateWhyMain failure mode to test
Public output changes only when content or code is releasedStatic generation plus CDN deliveryRemoves request-time rendering from the pathStale output after publish or incomplete rebuild
Public or semi-static HTML is identical for many visitorsShared response cachingReuses complete HTML and reduces repeated origin workWrong cache key, stale content, or low hit rate
Small request-specific decision can run without a distant data waterfallEdge computeMoves bounded logic closer to the requestRemote database/API calls erase the distance benefit
Authenticated or personalized work depends on data in one regionRegional application compute near the dataKeeps private logic and authoritative data close togetherLong visitor-to-region latency or serial dependencies
TTFB is dominated by queries, rendering, or API chains on every pathOrigin and dependency repairRemoves work that caching would only hideFix helps one route but leaves a shared dependency slow
TTFB is acceptable but LCP remains poorFrontend delivery and rendering repairTargets discovery, download, or render delayAn architecture migration changes nothing users perceive

Static generation is often the simplest option for content that changes through a controlled publishing event. Shared response caching suits content with an explicit freshness window and a safe invalidation path. Edge compute earns its place when request-time logic is bounded and its dependencies are local enough to preserve the latency gain.

Do not move an application to a distributed runtime merely because the platform offers one. Measure database, API, compute, and transfer phases first. A faster function start does not compensate for multiple cross-region backend calls.

For broader rendering and ownership trade-offs, use the headless web architecture guide and the Headless React versus monolithic CMS decision guide. Neither architecture label guarantees field performance.

What is safe to cache at the edge?

A shared cache is a data-sharing boundary. Before optimizing its hit rate, prove that two requests mapped to the same cache key are allowed to receive the same response.

Route classShared-cache starting positionRequired key or policy workRelease evidence
Public, versioned static assetCache with a content-hashed URL and long freshnessURL must change when bytes changeNew release serves new URL; old URL remains immutable
Public, identical HTMLEligible for shared cachingDefine query-string handling, host, path, and accepted encodingsHit and miss return equivalent correct content
Public, semi-static HTMLEligible with a freshness and purge policyDefine TTL, revalidation, publish event, and rollbackPublish, purge, stale, and origin-failure tests pass
Localized public routeEligible when locale is explicit and completePrefer a locale-specific URL; otherwise include every response-changing dimensionCorrect language, canonical, and alternate links in every state
Consent-dependent responseBypass unless variants are deliberately modelledConsent state must not leak or collapse into the wrong variantEvery consent state returns the correct scripts and content
Authenticated or user-specific responseDo not place in a shared cache by defaultUse private or no-store; design any exception as a security projectCross-account, logout, revocation, and failure tests pass
Personalized API responseDo not share by defaultBind access control to the data request, not only the page shellNo response can be replayed across identities

HTTP cache directives are easy to misread. MDN's `Cache-Control` reference explains that s-maxage sets freshness for shared caches, no-cache allows storage but requires validation before reuse, and no-store tells caches not to store the response. stale-while-revalidate permits reuse of a stale response during background revalidation; whether that is acceptable is a content and correctness decision.

Provider defaults differ. Cloudflare's default cache behavior currently says its CDN does not cache HTML or JSON by default and documents conditions that bypass caching. Vercel documents its own CDN cache and Cache-Control behavior. Treat both as platform-specific examples. Re-check the provider selected for the project instead of copying a header from a generic article.

Caching authenticated or personalized output incorrectly is a confidentiality incident, not a minor performance regression. Review cache-key inputs, cookies, query parameters, path normalization, redirects, and error responses. Include web-cache poisoning and cache-deception scenarios in the security review; Cloudflare's cache-security index is one vendor-specific starting point.

Instrument cache state and server work

A performance change cannot be evaluated if a response's cache state is unknown. Record a vendor-neutral internal label such as HIT, MISS, BYPASS, or STALE, and map it from the actual provider response headers. Keep the raw header for debugging, but normalize it for analysis.

Use Server-Timing to expose coarse request phases when the privacy and security review allows it. MDN's `Server-Timing` reference warns that the header may expose sensitive application or infrastructure information. Use short, stable categories; do not reveal table names, hosts, tenant identifiers, internal URLs, or query text.

The following TypeScript is illustrative. timings must come from the application's real instrumentation. Read the final cache status from the deployed cache layer's response headers; application code cannot know in advance whether a shared cache will serve a later request.

TYPESCRIPT
type SafeTiming = {
  appMs: number;
  dataMs: number;
};

export function addPerformanceHeaders(
  response: Response,
  timings: SafeTiming,
  cacheable: boolean,
) {
  const headers = new Headers(response.headers);

  headers.set(
    'Cache-Control',
    cacheable
      ? 'public, max-age=0, s-maxage=300, stale-while-revalidate=60'
      : 'private, no-store',
  );
  headers.set(
    'Server-Timing',
    `app;dur=${timings.appMs.toFixed(1)}, data;dur=${timings.dataMs.toFixed(1)}`,
  );
  return new Response(response.body, {
    status: response.status,
    statusText: response.statusText,
    headers,
  });
}

Do not use these TTLs as defaults. They exist only to show where a reviewed policy would be applied. Some providers consume, rewrite, or ignore directives; some frameworks own cache behavior above the raw response layer.

Inspect a response without downloading its body:

BASH
curl --silent --show-error --dump-header - --output /dev/null \
  https://www.example.com/representative-route

Repeat the request only when doing so matches the test design. A second request may be a hit, but a different point of presence, query string, cookie, deployment, or cache rule can create a different cohort. Save the request conditions with the headers.

For field attribution, Google's `web-vitals` library supports LCP, INP, CLS, and an attribution build with diagnostic fields. Record the route template, release, device class, market, cache state, and LCP subparts where your privacy policy permits it. Avoid sending full URLs or selectors when they may contain personal data.

The Edge Performance Proof Protocol

The protocol is designed to reject weak evidence before a rollout is declared successful.

1. Define the population and invariant

Choose representative routes and write down what must remain correct: status, canonical URL, language, consent behavior, identity boundary, price or availability freshness, structured data, and critical content. Define the target markets, devices, network profiles, and release identifier.

Do not mix a cached marketing page with an authenticated application route. Do not average a fast homepage with slow product, article, or checkout templates.

2. Capture the baseline

Record field TTFB and LCP by template, device, and market. Capture p50, p75, and p95 rather than the fastest observation. Note whether the value is URL-, group-, or origin-level and record the collection window and sample size.

Use lab traces to reproduce the field-attributed cause. The purpose of a lab run is diagnosis and regression testing, not replacement of field evidence.

3. Exercise every delivery state

Test the same routes and inputs in these states where they apply:

  • HIT: a reusable response is served from the shared cache;
  • MISS: the cache must fetch or generate the response;
  • BYPASS: policy sends the request directly to application or origin;
  • STALE: an allowed older response is served while revalidation occurs;
  • cold compute: the selected runtime has no warm execution context;
  • origin or dependency failure: stale-on-error, error, and recovery behavior are observed;
  • purge and publish: updated content becomes visible within the agreed freshness window.

Compare like with like. A warm hit after several test requests must not be compared with a cold baseline if ordinary users rarely reach that state.

4. Validate the complete LCP path

Check whether TTFB fell and whether field LCP moved with it. Break LCP into TTFB, resource load delay, resource load duration, and element render delay. If the remaining delay grows, the edge change may have shifted the bottleneck rather than improving the experience.

Monitor INP and CLS as guardrails. Faster delivery can change when JavaScript, fonts, consent UI, experiments, or recommendations execute. Edge caching does not directly remove long main-thread tasks or reserve layout space.

The Core Web Vitals and AI search guide contains the full field-to-fix workflow for LCP, INP, and CLS. This article remains focused on delivery architecture and the TTFB-to-LCP path.

5. Test correctness, security, resilience, and cost

Performance is one release dimension. Verify:

  • identity and authorization boundaries across accounts;
  • locale, currency, market, consent, and experiment variants;
  • cache invalidation after publish, price, inventory, or policy changes;
  • status codes, redirects, canonicals, robots directives, and structured data;
  • rollback and purge permissions;
  • origin load during misses and cache stampedes;
  • runtime, bandwidth, purge, logging, and observability costs;
  • behavior during cache-provider, application, and dependency failures.

A change that lowers TTFB but serves the wrong language, stale price, private account data, or an uncacheable error is a failed release.

6. Decide using pre-agreed evidence

Keep the change when the target field cohort improves and all correctness and reliability gates pass. Revise it when a specific state, such as misses in one region, fails. Revert it when safety, freshness, or reliability cannot be demonstrated.

Report technical and business outcomes separately. Lower TTFB or LCP does not prove a ranking, citation, conversion, or revenue change. Those outcomes need their own observation windows, controls, and limitations.

Why might a CDN leave Core Web Vitals unchanged?

A CDN can improve delivery without changing the dominant user-experience bottleneck.

  • Late LCP discovery: the hero image appears only after CSS or JavaScript runs.
  • Slow LCP transfer: the selected image is oversized or served from another slow origin.
  • Render delay: CSS, fonts, hydration, experiments, or long tasks prevent paint.
  • Poor INP: event handlers, third parties, or large DOM updates block the main thread.
  • Poor CLS: media lacks dimensions or late UI changes the layout.
  • Low cache eligibility: most responses are authenticated, personalized, or intentionally fresh.
  • Low hit rate: fragmented keys, rare routes, frequent purges, or region-specific traffic prevent reuse.
  • Remote dependencies: edge compute waits on an origin database or serial APIs.

Fix the dominant cause. If TTFB is already a small part of LCP, an edge migration is unlikely to be the first investment. If every miss exposes slow database or API work, repair that path even if caching hides it for popular routes.

Release guardrails before production

Use this final checklist:

  • [ ] The route is classified by data sensitivity, personalization, and freshness.
  • [ ] Every response-changing input is represented in the cache policy or causes a bypass.
  • [ ] Hit, miss, bypass, stale, cold, purge, and failure states are observable.
  • [ ] Server-Timing exposes no sensitive infrastructure or tenant details.
  • [ ] Baseline and post-release cohorts use the same routes, markets, devices, and definitions.
  • [ ] TTFB and all four LCP subparts are reviewed together.
  • [ ] INP and CLS remain within agreed regression limits.
  • [ ] Locale, consent, identity, canonical, status, and structured-data checks pass.
  • [ ] Invalidation, rollback, and incident owners are documented.
  • [ ] Cost and origin-load changes are reported beside latency changes.
  • [ ] Search, citation, engagement, and conversion outcomes are reported separately.

If the team cannot complete the cache-safety matrix, begin with measurement and origin diagnosis. If it can prove a cacheable response and an attributed request-path bottleneck, pilot one representative route before expanding the policy.

AppWebSeo's Core Web Vitals audit identifies whether the dominant delay is in the request path, critical-resource delivery, rendering, interaction, or layout. When evidence points to server and network latency, the edge infrastructure review can define route-level caching, regional delivery, instrumentation, and release gates. Either review may conclude that origin repair or a frontend fix is safer than an edge migration.

Sources, verification, and limitations

Primary documentation was checked on 2026-09-20:

Platform defaults, headers, runtimes, and cache behavior can change. Re-check the chosen provider during implementation. Chrome UX Report may not have URL-level data for every route, and aggregated data may not isolate a release quickly enough; first-party real-user monitoring requires its own privacy, sampling, retention, and data-quality review.

The equations and Edge Performance Proof Protocol are diagnostic planning tools. They are not measured results for AppWebSeo or a customer, and they do not guarantee a latency, Core Web Vitals status, ranking, AI citation, conversion, or revenue outcome. The legacy URL is retained for continuity; “sub-second” is not the promise of this article.

A

AppWebSeo

SEO & Engineering Editorial Team

Specializing in high-performance web systems, Generative Engine Optimization, and enterprise AI architecture at AppWebSeo.

Share this Technical Breakdown

Forward this architecture guide to your team, colleagues, or engineering network.

Transform These Insights into Production Architecture

Schedule a technical architecture review with our senior engineering team.

All Topics