Video Call API Pricing: Cost and Usage Model Explained

TL;DR: Most video call APIs bill participant minutes, not sessions. A one hour call with ten people is 600 billable minutes, not 60. That single misunderstanding is behind most surprise invoices. Recording storage, resolution multipliers, and retention add more on top. Model your real usage before you compare headline rates, because the rate is the smallest part of the answer.

A product team budgets for video by taking the vendor’s per-minute rate and multiplying it by expected meeting hours. The number looks comfortable. They ship.

The first real invoice is roughly eight times the estimate. Nothing went wrong and the vendor did not overcharge. The team modelled sessions when the vendor bills participants, and every meeting had about eight people in it.

This is the most common costing mistake in video, and it is entirely avoidable. This guide explains how video call API pricing actually works, the four billing models you will encounter, the costs that never appear in the headline rate, and how to forecast your bill before you commit.

Most video call APIs multiply participants by minutes by a rate. Ten people in a one hour meeting produce 600 participant minutes, not 60. Recording, transcription, and higher resolutions are usually billed separately on top of that base figure.

The industry standard formula is simple once you see it: participants multiplied by minutes multiplied by rate. Published analysis of video API pricing models puts the typical rate near 0.004 US dollars per video minute, though this varies by vendor and resolution.

The trap is that costs scale with group size, and group size is the variable product teams control least. As one analysis of usage-based pricing notes, many founders assume they pay per session when they actually pay per participant per minute. For group-based products, costs climb far faster than expected.

A worked example makes the scale obvious. These figures are illustrative, using the published 0.004 rate.

ScenarioSessions/monthAvg participantsParticipant minutes
1:1 support calls, 30 min500230,000
Team meetings, 60 min5008240,000
Online classes, 90 min20040720,000

Same vendor, same rate, twenty four times the cost between the first and third row. Session count barely moved. Participant count did all the work.

The four pricing models you will encounter

Vendors rarely describe their model in these terms, so it is worth naming them yourself during evaluation.

ModelYou pay forPredictable?Best fit
Participant minutesEvery person, every minuteLowSmall calls, spiky usage
Monthly active usersUnique users per monthMediumHeavy use by a stable base
Seats or plansLicensed hostsHighInternal use, known headcount
Self-hostedInfrastructure and staffHigh, but fixedSovereignty needs, large scale

Participant-minute billing punishes exactly the thing many products want, which is large engaged sessions. Seat and plan pricing decouples cost from session size entirely, which is why platforms serving classrooms and all-hands meetings usually price that way. Convay is in that group, selling plans rather than metered participant minutes, so a session with 300 attendees does not cost thirty times a session with ten.

Neither model is better in the abstract. They are better for different usage shapes, and the only way to know which suits you is to model both against your own numbers.

What costs are missing from the headline rate?

Recording is usually billed per recorded minute and then again for storage each month it is retained. Higher resolutions often carry multipliers. Transcription, live streaming, and phone dial-in are typically separate line items. Compliance retention quietly turns storage into a permanent growing cost.

The base rate is the beginning of the bill, not the end. These are the additions that catch teams out.

Retention deserves its own conversation with your compliance team before you sign anything. If your sector requires you to keep recordings for years, storage is not a small add-on. It becomes a permanent, compounding line item that grows every month you operate, and it is frequently larger than the call cost itself by year three.

How do you forecast your video API bill?

Estimate monthly sessions, average duration, and average participants, then multiply all three to get participant minutes. Add recorded minutes and multiply storage by your retention period. Model the same numbers at three times growth, because pricing models that work at launch often break at scale.

Build the forecast in five steps, and do it in a spreadsheet you can hand to finance.

  1. Count expected sessions per month.
  2. Multiply by average session length in minutes.
  3. Multiply by average participants. This is your participant-minute base.
  4. Add recording: what share of sessions get recorded, and what is the retention period in months.
  5. Repeat the whole calculation at three times your expected volume.

That last step matters most. Metered pricing is usually cheapest at low volume and most expensive at scale. Plan and seat pricing is the reverse. The crossover point decides which vendor is genuinely cheaper for you, and it rarely sits where the sales deck implies.

If you expect large sessions, check whether large-scale meetings are priced differently from standard ones. Many platforms treat them as a separate product with separate commercial terms.

When does self-hosting become cheaper?

Self-hosting removes per-minute costs entirely and replaces them with infrastructure and staff. The licence may be free. The operation is not.

Analysis of open-source video deployments makes this point directly: the cheapest tools by licence cost frequently become the most expensive once development and operations time is counted. Running production video means media servers, network traversal infrastructure, monitoring, scaling, and someone on call when a session fails during a board meeting.

Self-hosting usually wins in two situations. The first is very high volume, where per-minute charges exceed the fully loaded cost of running your own infrastructure. The second is regulatory, where data sovereignty requirements make external hosting unacceptable at any price. In the second case the cost comparison is not really the point, because the alternative is not permitted.

A middle path exists. On-premise deployment of a commercial platform gives you the sovereignty position without asking your team to become WebRTC infrastructure engineers.

Conclusion

Three things will keep your video budget honest.

Model participant minutes, never sessions. Group size drives your bill far more than meeting count does.

Price recording and retention as a separate long-term commitment, because storage you are required to keep never stops accruing.

Run the numbers at three times your expected volume. The cheapest vendor at launch is often the most expensive one the year after.

Frequently Asked Questions

Most vendors multiply participants by minutes by a rate. A one hour meeting with ten people produces 600 participant minutes, not 60. Recording, transcription, and higher resolutions are usually billed separately on top of that base figure.

A participant minute is one person spending one minute in a call. Four people in a fifteen minute meeting generate 60 participant minutes. It is the standard billing unit for metered video APIs, and confusing it with session minutes is the most common cause of underestimated video budgets.

The usual causes are modelling sessions instead of participant minutes, underestimating average group size, and overlooking recording storage that accrues every month a file is retained. Resolution multipliers and separately billed transcription also push the effective rate above the advertised one.

Usage-based pricing is generally cheaper at low volume and with small groups. Seat or plan pricing becomes cheaper as sessions get larger, because cost stops scaling with attendee count. Model both against your own forecast, then repeat at three times volume to find the crossover point.

Sometimes, but not as often as licence costs suggest. Self-hosting replaces per-minute charges with servers, bandwidth, and engineering time, and analysis of open-source deployments finds the cheapest licences often produce the highest total cost. It wins clearly at very high volume, or where data sovereignty rules make external hosting unacceptable.

Want a cost model built on your actual usage?

Bring your session count, average group size, and retention requirement. We will model it against plan pricing and show you where the crossover against metered billing sits for your volume.

Share the Post:

See Convay in action

Secure, AI-powered video conferencing built for enterprises and government organisations.

Contact Us
Picture of Fariduzzaman Swadhin

Fariduzzaman Swadhin

Fariduzzaman Swadhin is a professional in the tech industry, specifically known as a AI iSaaS Analyst Growth and Product Marketing Manager. He currently works at Convay, a secure collaboration platform, where he focuses on driving revenue and retention through Go-to-Market (GTM) strategies and Product-Led Growth.