Uncategorized

How to Benchmark Website Performance Before and After Adding an AI Feature

Black server cabinets with rack-mounted equipment and cables routed overhead.

Answer box: Compare the same page with the AI feature disabled and enabled, under matching browser, device, network and cache conditions. Measure page loading separately from chatbot response time, then repeat the tests and report the median and range—not just one performance score. Lighthouse results naturally vary, and a normal page-load audit does not measure an entire chat conversation. (github.com)

1. Separate page performance from AI response time

Your benchmark should answer three questions:

  • Does the feature affect visitors who never use it? Test initial loading with the widget enabled but untouched.
  • Does interacting with it make the page less responsive? Open the panel, type, submit a question and use other page controls.
  • How long does an answer take? Measure the interval from submission to the first visible answer content and to completion.

Keep those results separate. Interaction to Next Paint (INP) measures responsiveness to user input through the next paint; it is not a measure of the full wait for an asynchronous AI answer. A promptly displayed loading indicator and a slowly generated answer can therefore represent different performance problems. (web.dev)

Use these metrics as your starting point:

Measurement What it helps assess
Largest Contentful Paint (LCP) When the largest qualifying visible content appears
Cumulative Layout Shift (CLS) Unexpected movement of visible content
INP during interaction Responsiveness to clicks, taps and keyboard input
Total Blocking Time (TBT) Blocking during a Lighthouse loading test
Document Time to First Byte (TTFB) A supporting clue when investigating slow loading

Lighthouse’s normal navigation audit cannot measure INP without user interaction. TBT is a useful laboratory proxy, not an interchangeable INP result. (web.dev)

Google’s current “good” Core Web Vitals thresholds are LCP at or below 2.5 seconds, INP at or below 200 milliseconds, and CLS at or below 0.1. Assess these at the 75th percentile of real visits, separating mobile and desktop—not by treating a small laboratory sample as a field assessment. (web.dev)

2. Create a baseline you can reproduce

Before installing or changing a plugin, preserve a restorable backup and record the configuration you intend to restore. Start in a protected staging environment.

Choose a representative page where the feature will appear. For a site-wide widget, also include another important page template rather than assuming one page represents the entire website.

Create two explicit states:

  • A — Disabled: the feature’s integration is removed or genuinely switched off.
  • B — Enabled: the intended launch configuration is active.

Verify state A in the browser’s Network panel. Hiding the chat button is not a sufficient control: check whether its scripts, styles, frames or requests still appear. DevTools exposes request initiators, resource types, transfer sizes and timing details that help verify what actually loaded. (developer.chrome.com)

For WordPress, record whether you disabled the widget through a plugin setting or deactivated the entire plugin. Those comparisons answer different questions; use the setting that matches your intended deployment.

Write down:

  • Page URL and deployment identifier.
  • Plugin or integration version and settings.
  • Logged-in or logged-out state.
  • Consent choices and other enabled widgets.
  • Hosting environment and cache configuration.
  • Test date and tool versions.

Do not update the theme, optimize images or change hosting between A and B. Lighthouse’s own guidance identifies changing page content and underlying test conditions as sources of variability. (github.com)

Treat staging results as preliminary if its resources, CDN or cache behavior differ from production.

3. Hold device, network and cache conditions constant

Use the same test computer, browser version, viewport, network profile and CPU-throttling setting for each comparison. Avoid running tests simultaneously or while unrelated applications consume substantial resources; hardware, network conditions and resource contention can change Lighthouse results. (github.com)

Keep separate result sheets for mobile and desktop. Record the exact throttling settings rather than labeling a test only “mobile.”

Test two browser-cache conditions:

  • Cold browser cache: an initial visit without reusable browser assets.
  • Warm browser cache: a repeat visit with caching allowed.

Chrome documents that repeat visits can reuse cached resources. Its Empty Cache and Hard Reload workflow helps examine a first-visit loading experience. (developer.chrome.com)

Do not assume that clearing the browser cache also resets your CDN, server-side page cache or AI-answer cache. Record each layer’s state separately and mark anything you cannot verify as unknown.

A useful comparison matrix is:

Feature state Browser-cache state Activity
Disabled Cold and warm Load the page
Enabled Cold and warm Load without opening chat
Enabled Cold and warm Open chat and submit a fixed question

Alternate disabled and enabled runs rather than completing all baseline tests first. This is a practical test-design choice intended to reduce the influence of changes over time, not a guarantee that every source of noise disappears.

4. Inspect browser overhead, not just the overall score

Run Lighthouse for loading diagnostics, then use DevTools to investigate the difference.

In the Network panel, compare:

  • Added JavaScript, CSS, fonts and frames.
  • Transferred bytes and request counts.
  • Third-party domains.
  • Requests triggered during loading versus after opening chat.
  • Failed or repeatedly retried requests.

Use the Initiator information to connect a request to the integration that triggered it, and the Timing tab to examine individual requests. DevTools provides these details directly. (developer.chrome.com)

Next, record a Performance trace while opening the widget and submitting a question. Inspect scripting, rendering and main-thread activity. Chrome’s Performance panel is designed for runtime analysis, which complements a loading audit. (developer.chrome.com)

Include ordinary website actions in this session: open navigation, type into a form and click another control while the answer is being produced. Your test should assess the page around the chatbot, not only the chatbot itself.

Check for layout movement when the launcher or conversation panel appears. Save the trace or screenshots that show the suspected problem instead of relying solely on a changed score.

If you experiment with loading the widget later, test its first opening again. Treat reduced startup work and slower first-use availability as separate outcomes to evaluate.

5. Measure the AI request path independently

For answer generation, define these milestones before collecting timings:

  1. The user submits the question.
  2. The interface acknowledges submission.
  3. The first actual answer text becomes visible.
  4. The answer finishes, or an error is shown.

A spinner is acknowledgement, not answer content. Likewise, an HTTP response beginning is not necessarily the first visible text.

Developers can use named performance.mark() entries and performance.measure() durations for application-specific milestones. These integrate with browser performance tooling. Place the markers at the real application events; loading an isolated timing snippet would not measure the conversation. (developer.mozilla.org)

Where backend instrumentation is available, record separate durations for:

  • Queue waiting.
  • Application preparation.
  • Retrieval or database work.
  • The upstream AI request.
  • Delivery and rendering of the response.

Use a correlation identifier to connect the request’s records. OWASP describes interaction identifiers as a way to link related application events and recognizes application logging as useful for performance monitoring. (cheatsheetseries.owasp.org)

Use fixed, non-sensitive test questions and consistent conversation history. Record the model identifier, output limit and streaming setting where the integration exposes them.

If a closed plugin provides no backend timings, report only the end-to-end measurements you can observe. Do not attribute the delay specifically to hosting, retrieval or the AI provider without supporting evidence.

Exclude credentials and sensitive conversation content from benchmark logs and shared exports. OWASP’s logging guidance specifically identifies access tokens, passwords, primary secrets and sensitive personal data as information requiring exclusion or protective handling. (cheatsheetseries.owasp.org)

Black server cabinets with rack-mounted equipment and cables routed overhead.
Server racks are an infrastructure illustration, not evidence about the hosting used by a particular chatbot. — Jemimus. https://www.flickr.com/photos/12967790@N00/2528960658/ Source CC BY 2.0

6. Label cold starts, caching and repeat-run variability

A cold browser cache and a cold backend start are different conditions.

For example, AWS Lambda documents that preparing a new execution environment adds initialization latency, while a later invocation may reuse an existing environment. That establishes a possible explanation for some serverless deployments—not proof that your plugin has the same behavior. (docs.aws.amazon.com)

Keep the first request after inactivity separate from immediate follow-up requests. Call it a confirmed cold start only when platform telemetry supports that label.

Also distinguish a cached answer from a newly generated answer. If your integration does not expose cache status, mark it unknown rather than inferring it solely from a fast response.

For an initial investigation, choose a fixed number of runs—such as five per condition—and increase it when results are inconsistent. This is a practical starting point, not a statistically validated sample-size requirement.

Report the run count, median, minimum and maximum. Keep errors and timeouts visible rather than calculating an apparently fast result from successful requests alone.

Lighthouse recommends repeated testing because performance scores can change even without a code change. Small differences within an unstable set of results should remain inconclusive until further testing supports them. (github.com)

7. Work through a comparison without inventing results

Consider a blog adding a support chatbot to an article page. The following is a worked testing procedure, not a claim about measured performance.

First, collect disabled and enabled runs under the same cold-browser-cache profile. Leave the chatbot untouched during these loading tests.

Then summarize your recorded values:

Comparison Calculation
LCP change Enabled median LCP − disabled median LCP
LCP percentage change LCP change ÷ disabled median LCP × 100
Added transfer Enabled transferred bytes − disabled transferred bytes
Response-generation interval Completion timestamp − first-answer timestamp

Next, run an interaction session: open chat, submit the fixed question, and click the site navigation while the answer arrives. Capture the trace and the application’s timing milestones. User Timing supports the custom intervals, while the Performance panel supports runtime investigation. (developer.mozilla.org)

Interpret the evidence conditionally:

  • If enabled runs repeatedly load more slowly, investigate the added startup work.
  • If loading remains similar but answer delivery is slow, investigate the request path.
  • If other controls become sluggish during chat activity, inspect the runtime trace.
  • If results overlap substantially and vary widely, report an inconclusive comparison.

Your report should state the page, configuration, conditions, sample count, observed differences and limitations. Do not replace those details with “the plugin is fast.”

8. Check common mistakes before deciding to launch

Before accepting the feature, check that you have not:

  • Used one Lighthouse score as the entire benchmark.
  • Hidden the launcher while leaving its integration active.
  • Compared different devices or throttling profiles.
  • Mixed first visits with repeat visits.
  • Treated TBT as measured INP.
  • Counted a spinner as the first answer.
  • Removed failed requests from the report.
  • Assumed staging proves production behavior.

After a controlled release, monitor real-user experience as well as laboratory results. PageSpeed Insights field data covers a trailing 28-day collection period; it is not an immediate before-and-after reading. Check whether it reports the specific URL or falls back to origin-wide data. A missing field report does not establish good performance. (developers.google.com)

Keep a documented rollback path and define acceptable regressions before reviewing results. If a material problem appears, disable the integration, verify that its resources stop loading, and repeat the baseline checks.

The defensible conclusion is limited but useful: whether this configuration produced a repeatable difference on the tested pages under the recorded conditions.

Read next: How to Stop a Staging Website from Sending Emails, Taking Payments or Running AI Jobs