Andrew Crossley
    Back to home

    CASE STUDY • AI • VOICE • VERTICAL SAAS

    CallFlow AI

    An AI receptionist designed for trades, clinics, salons and appointment-led businesses. It answers calls, identifies customer intent, qualifies enquiries and moves suitable customers toward an appointment or callback.

    STAGE
    Live product prototype
    SECTOR
    Appointment-led small business services
    MY ROLE
    Product creator & product lead
    LIVE PRODUCT
    call-flow.co.uk

    The problem

    An appointment-led business earns from the phone and cannot answer it. The plumber is under a sink, the clinician is with a patient, the stylist is mid-appointment. The call goes to voicemail and the caller rings the next business on the list.

    Missed calls are not a communications problem, they are a revenue problem. But hiring a receptionist for a two-person business does not add up, and generic call-answering services do not know enough about the job to be useful.

    Who the customer is

    Small, appointment-led service businesses: trades, clinics, salons, garages, veterinary practices. Typically one to fifteen people, no dedicated front desk, and a diary that is the core operational asset.

    What users currently do instead

    • Voicemail, which most callers will not use.
    • Call diversion to a mobile that is already busy.
    • A shared human answering service that takes a message but cannot answer whether you cover that postcode or fit that boiler.
    • Returning missed calls in the evening, by which point the customer has booked elsewhere.

    The core product journey

    • A call comes in and the AI answers in the business's own framing: what it does, where it covers, what it can book.
    • It identifies intent — new enquiry, existing customer, quote request, complaint, supplier, nuisance call.
    • For a bookable enquiry it qualifies the essentials: job type, location, urgency, contact details.
    • It offers real appointment slots and confirms one, or captures a callback where booking is not appropriate.
    • Anything outside its competence is escalated to a human rather than guessed at.
    • The owner gets a transcript, a structured summary and a qualified lead record instead of a missed-call notification.

    The hardest product decision

    How much autonomy to give the AI. A receptionist that only takes messages is not worth paying for. A receptionist that confidently books the wrong job into a real diary is worse than voicemail.

    The decision was to make escalation a first-class product feature rather than an error state, and to design the evaluation before the personality. The system is judged on whether it captured the right facts and made a safe booking decision, not on whether it sounded impressive.

    What I deliberately did not build

    • Outbound calling and sales dialling. Different product, different regulatory surface, different buyer.
    • Full CRM functionality. It captures qualified leads and hands them on rather than trying to become the system of record.
    • Unrestricted open-ended conversation. The AI operates inside a defined scope and escalates outside it.
    • Multi-language support, until there is demand evidence from real users rather than an assumption.

    Business model

    Subscription per business, with usage sitting alongside it, because voice minutes and model calls are a real variable cost that has to be modelled rather than hoped away.

    Pricing is designed against the value of a recovered booking for that trade, not against a competitor's price list. This is a hypothesis that needs paid customers to confirm.

    Technology and AI approach

    Speech-to-text, a language model constrained by a business-specific brief, and text-to-speech, wrapped in explicit conversation state so the call cannot wander.

    Calendar integration for real availability rather than a promise to call back.

    Evaluation is designed into the product: transcripts are scored on whether intent was correctly identified, whether required fields were captured, whether the booking was safe, and whether escalation happened when it should have. Latency is treated as a product requirement because a pause on a phone call reads as a hang-up.

    Screenshots

    The product surfaces that carry the journey.

    Incoming call

    CallFlow AI incoming call screen showing a live customer call being answered by the AI receptionist

    Incoming call

    AI qualification

    CallFlow AI qualification view showing captured customer intent and enquiry details

    AI qualification

    Calendar booking

    CallFlow AI calendar booking screen offering available appointment slots to a caller

    Calendar booking

    Call transcript

    CallFlow AI call transcript showing the full conversation between caller and AI receptionist

    Call transcript

    Human escalation

    CallFlow AI human escalation screen handing a complex call to a member of staff

    Human escalation

    Lead dashboard

    CallFlow AI lead dashboard listing qualified enquiries captured from inbound calls

    Lead dashboard

    What this demonstrates

    • Building a voice AI product where the failure states matter more than the demo.
    • Designing evaluation before launch rather than discovering reliability problems in production.
    • Scoping AI autonomy deliberately, with escalation as a designed path.
    • Pricing an AI product with variable inference cost built into the model.

    What has been validated

    • The end-to-end journey works as a live prototype: answer, qualify, book, transcribe, escalate.
    • The missed-call problem is well established across appointment-led trades.

    What still needs validation

    • Willingness to pay, and at what price point, for a specific trade.
    • Real-world accuracy across accents, noisy environments and interrupted callers at volume.
    • Whether business owners trust an AI with their live diary without a supervision period first.

    What I would test next

    • A small number of paid pilots in one trade rather than several, so the evaluation set is comparable.
    • A measured qualification-accuracy rate against human review of the same transcripts.
    • A test of whether owners keep autonomous booking switched on after the first fortnight.