Uncategorized

What to Monitor When Your Website Depends on an AI API

Answer box: Monitor whether visitors can finish the AI task—not just whether the page loads. Track page availability, backend errors, provider latency, queue waiting time and feature completion separately. Combine real-request metrics with a small scheduled test through your website. Alert on sustained visitor-facing failures, and test a clear fallback in staging before relying on it. External checks show the visitor’s experience; internal metrics help explain the cause. (grafana.com)

1. Define what “working” means for your AI feature

For an AI feature, distinguish page available, request accepted, provider response received and result delivered. Treat them as separate checkpoints. Prometheus recommends monitoring online services from both client and server perspectives because differences between them help reveal problems. (prometheus.io)

For a chatbot, write an explicit completion rule:

  • The website accepts a valid message.
  • The backend starts the intended workflow.
  • The provider returns a response, or an approved fallback is selected.
  • The browser displays the result and leaves its loading state.
  • The workflow finishes within your chosen deadline.

Record an AI answer and a fallback as different outcomes. A help-page link may keep the site useful during an outage, but it is not a completed AI answer.

Decide how to classify expected refusals, invalid inputs and visitor cancellations. Keep those visible without automatically calling each one an infrastructure failure. Mark synthetic requests separately from visitor traffic.

Availability also does not prove answer quality. OpenAI documents that model outputs are variable and recommends application evaluations to check behavior. A chatbot that returns text reliably may still return unsuitable text. (developers.openai.com)

2. Track five signals instead of one uptime percentage

Use this as a proposed dashboard design. Each row answers a different operational question.

Signal What to record Question it answers
Page availability External page-check success and duration Can a visitor reach the site?
Backend errors Application failures, timeouts and provider error categories Where are requests failing?
API latency Provider-call duration; for streaming, first-content and full-response timing Is the AI dependency responding promptly?
Queue waiting time Queue depth, oldest pending job age and last successful worker activity Is accepted work actually moving?
Feature completion Completed AI results divided by eligible attempts, plus fallback and cancellation counts Are visitors receiving the intended result?

These choices build on Prometheus guidance to measure requests, errors, latency, in-progress work and queue waiting time. (prometheus.io)

Measure end-to-end duration as well as provider duration. Instrument the milestones from submission to display so you can investigate a gap between them. For streaming chat, make the first visible content and final completion separate milestones rather than treating the start of a stream as success.

For queued work, prefer an age check alongside a depth check. Define age as the time since the oldest pending job was accepted. If the queue is empty, display “empty,” not a missing-data failure.

Classify provider failures by status and structured error code. OpenAI’s documentation distinguishes authentication failures, rate limiting, exhausted credits or limits, and server errors. In particular, a 429 does not identify one universal cause or recovery action. (developers.openai.com)

3. Use layered synthetic checks through your website

A synthetic check is a scheduled test with a defined success condition. Grafana’s documentation describes both HTTP checks and scripted browser workflows, complementing metrics collected inside the application. (grafana.com)

For a small chatbot, consider three layers:

  1. Page check: Confirm that the public page responds and includes a stable expected element.
  2. Workflow check with a stub: Exercise the interface and backend using a controlled replacement for the provider.
  3. Live-provider check: Send a small request through the actual application workflow and verify completion.

The stubbed check is useful for testing application behavior without making a live model call. Keep its results clearly labeled: passing a stubbed test does not establish provider availability.

For the live check, use harmless synthetic input, a restricted test identity and the production model or route you actually depend on. Request a brief answer and set a modest output limit supported by that endpoint. Avoid private documents, visitor conversations and external actions.

Assert the outcome, not just the HTTP status. For example, require a nonempty answer, a completed workflow state and a cleared loading indicator. Grafana’s browser-check tutorial uses assertions to validate page elements and workflow results rather than assuming requests succeeded. (grafana.com)

Avoid brittle exact-sentence matching. Separate basic completion checks from a broader quality-evaluation suite.

Finally, budget for probe count. In Grafana Synthetic Monitoring, every selected probe executes independently at the scheduled interval; additional locations multiply executions. (grafana.com)

4. Build a small monitoring plan with actionable alerts

The following is an illustrative plan, not a benchmark, measured result or provider recommendation. Assume a low-traffic chatbot with an application-chosen 20-second answer deadline.

Check or metric Proposed starting configuration Proposed action
Public page One external probe every minute Notify after three consecutive failures
Live chatbot workflow One probe every 15 minutes Investigate after two consecutive failures
Visitor completion Rolling 10-minute window Notify below 95%, only with at least 20 eligible attempts
Queue age, if applicable Evaluate every minute Warn when oldest pending work exceeds 10 seconds for three minutes
Authentication or quota failure Classify each failure Send a configuration/billing notification; do not repeatedly retry
Monitoring freshness Track expected check executions Notify when scheduled checks stop reporting

At one live request every 15 minutes, the planned volume is:

4 requests per hour × 24 hours = 96 requests per day.

That assumes one probe, one provider call per execution and no retries. Extra tool calls, retries or probe locations increase it. Estimate spending from your current model prices, measured token usage and monitoring-service charges; a short prompt alone does not establish a dollar cost.

In a hypothetical window with 40 eligible attempts and 37 completed AI answers:

37 ÷ 40 × 100 = 92.5% completion.

That crosses the example threshold. Record any fallbacks separately so they do not conceal the decline.

For a quieter site, the traffic threshold may never be reached. Retain the synthetic alert rather than treating insufficient traffic as evidence of health. Two consecutive failures on a 15-minute schedule can take roughly 15–30 minutes to detect, depending on when failure begins.

Tune thresholds to your baseline and visitor deadline. Prometheus recommends alerting on user-facing symptoms, allowing room for brief blips and avoiding notifications with no useful action. (prometheus.io)

Give every alert an owner, dashboard reference and first action. Test delivery to the person responsible, not just the rule’s status.

5. Keep diagnostic logs useful without storing conversations

OpenAI recommends retaining provider request IDs for troubleshooting. Its reference identifies x-request-id and rate-limit response headers as diagnostic information. (developers.openai.com)

A proposed allowlisted log record could include:

  • Timestamp and environment.
  • Internal correlation ID.
  • Feature and deployment version.
  • Provider request ID, when available.
  • Error category and structured error code.
  • Queue, provider and total durations.
  • Retry count.
  • AI-completion or fallback outcome.
  • Synthetic-test flag.

Do not put prompts, responses, authorization headers, session cookies or retrieved document contents in routine monitoring records. Prefer selecting safe fields over collecting whole payloads and trying to clean them afterward. OpenTelemetry’s redaction guidance supports modifying sensitive log attributes before export. (opentelemetry.io)

Run a test with obviously fake sensitive values, then inspect application logs, exported telemetry and alert messages. Verify redaction at every destination you use, including exception reporting.

Keep unique request IDs in restricted logs or traces rather than metric labels. Prometheus warns that label combinations create additional time series and advises against high-cardinality labels. (prometheus.io)

6. Design fallback behavior and bounded retries

Choose fallback behavior before an incident. For a chatbot, a proposed fallback might stop the loading indicator and say:

“The assistant is temporarily unavailable. You can still browse the help pages. Please try again later.”

Offer only alternatives that actually exist. Do not promise human follow-up unless a staffed process supports it.

Keep the ordinary website usable where your architecture permits. Record fallback activation as a degraded outcome, and monitor it separately from successful AI answers.

For retryable failures, OpenAI’s rate-limit guidance recommends honoring a valid Retry-After header or using exponential backoff with jitter. It also advises limiting attempts and total retry time, accounting for SDK retries, and not retrying billing or quota errors that require action. Unsuccessful requests contribute to per-minute limits. (developers.openai.com)

Set the combined queue, request and retry budget within your visitor-facing deadline. Verify SDK defaults instead of adding another retry loop blindly.

Use the provider status page as supporting evidence, not your sole monitor. OpenAI states that its availability metrics are aggregated across tiers, models and error types; an individual customer’s experience can differ. (status.openai.com)

7. Simulate provider failure safely, then verify recovery

Use this proposed drill in an isolated staging environment. Maintain a known-good deployment and rollback path before changing test configuration.

  1. Isolate the environment. Use synthetic data and separate queues and credentials. Disable payments, outgoing email and unrelated automation.
  2. Substitute a test provider adapter. Make it return a simulated server error, rate-limit error, malformed result or delayed response without contacting the provider.
  3. Exercise the browser workflow. Submit a harmless message through the same interface a visitor uses.
  4. Verify behavior. Check the deadline, retry limit, cleared loading indicator, truthful fallback and absence of duplicate work.
  5. Verify monitoring. Confirm the right error category, failed AI completion, fallback event, queue state and notification.
  6. Restore and retest. Remove the simulation, verify configuration, run a normal check and confirm recovery is reported.

For streamed chat, add a case where a response starts and then stops before completion. For queued features, simulate a stalled staging worker and verify that age monitoring detects it.

Do not force a real outage by exhausting quota, revoking production credentials or blocking production network access. Keep failure injection unavailable to ordinary visitors and disabled in the production configuration.

Before closing the drill, check three common mistakes: Did the monitor accept a fallback as AI success? Did the test bypass the actual interface? Did the dashboard report healthy when monitoring data was missing?

Also test the notification path itself. Prometheus explicitly recommends monitoring the monitoring system and checking that alerts reach their destination. (prometheus.io)

Read next: How to Stop a Staging Website from Sending Emails, Taking Payments or Running AI Jobs