Paul S.

Warranty Checkout Migration

Self-service warranties were the first checkout channel moved onto the new platform: the narrowest funnel that still exercised the whole path. The migrated flow generated $500K+ in its first two months and took end-to-end checkout from ~10–15s to under 2s.

Context

Self-service warranties ran on the legacy stack with limited operational visibility. The revenue was real but small next to the main hardware funnels, which is what made it a workable place to run the new platform against production traffic.

Problem

The cart-checkout platform was built and passing its own tests, but nothing had run on it in production. We needed a funnel carrying real orders, moved across without disrupting customers or losing revenue, and we needed to find out how the platform behaved on call before anything larger depended on it.

What I did

  • Modeled the warranty checkout flow on the new platform, including pricing rules, eligibility, and downstream order workflows.
  • Built a staged cutover: dual-running, percentage ramp, kill-switch, and clear rollback steps.
  • Designed dashboards and alerts specific to the cutover, covering error taxonomy, latency percentiles, and revenue parity.
  • Worked with product, support, and partner engineering teams to handle edge cases before they reached customers.

Choosing the first funnel

A warranty order touches catalog, cart, pricing and eligibility rules, tax, payment authorize/capture, order creation, and downstream fulfillment workflows. Every component of the platform runs on one, so the funnel exercised the whole path while carrying a small fraction of the revenue the hardware funnels carry.

Low-traffic countries narrowed the audience on the same principle. We widened the funnel and the geography one step at a time, so that when a number moved we knew which change had moved it.

Rollout

The migrated flow ran in parallel with legacy and took traffic by percentage. Every step was reversible from the router, without a code change or a deploy.

  1. Dual-run: both flows live, legacy still authoritative, the new flow observed but not trusted.
  2. Ramp: a small traffic percentage moves to the migrated flow behind the router.
  3. Hold and compare: guardrails evaluated against the legacy cohort, not against an absolute target.
  4. Widen, or roll back: increase the percentage, or flip the kill switch and return to legacy in seconds.

The router evaluated the same four guardrails it used for every funnel: conversion, payment success, p95 latency, and error budget. Each was compared against the traffic still sitting on legacy at that moment. A fixed p95 target would have failed the migration on any day the whole site was slow, for a problem the migration did not cause.

fun goNoGo(migrated: Cohort, legacy: Cohort): Decision = when {
    migrated.conversion    < legacy.conversion.withTolerance()    -> RollBack
    migrated.paymentSuccess < legacy.paymentSuccess.withTolerance() -> RollBack
    migrated.p95           > legacy.p95.withTolerance()            -> Hold
    migrated.errorBudgetBurn > BUDGET_THRESHOLD                    -> Hold
    else -> Widen
}
Every guardrail is evaluated as migrated-vs-legacy on concurrent traffic, so a bad day for the whole site doesn't read as a failed migration.

What the ramp measured

A migrated checkout can be fast, return 200s, and still lose money. An eligibility rule that rejects slightly too often, a payment method that stops being offered, a tax call that fails in one market only: none of those show up in latency or error rates.

So revenue parity was tracked per ramp step alongside the technical guardrails. The error taxonomy separated customer-caused failures from system-caused ones, because a declined card and a failed tax call both end as 'checkout didn't complete', and counting them together would have hidden a broken tax integration behind normal card declines.

The dual-run, ramp, comparative-guardrail, kill-switch sequence became the template the next funnels followed.

Result

  • $500K+ in revenue in the first two months on the migrated flow, the first production evidence that the platform could carry orders end to end.
  • End-to-end checkout ~10–15s → under 2s, by replacing per-surface legacy integrations with a single orchestration call path.
  • A migration template (dual-run, ramp, comparative guardrails, kill switch) that every later funnel followed.
  • The platform ran on production traffic before any hardware funnel depended on it.

Tech

  • Kotlin
  • GraphQL
  • Commercetools
  • Stripe
  • Affirm
  • AWS
  • EKS
  • Datadog
  • PagerDuty