TL;DR: Most video call APIs bill participant minutes, not sessions. A one hour call with ten people is 600 billable minutes, not 60. That single misunderstanding is behind most surprise invoices. Recording storage, resolution multipliers, and retention add more on top. Model your real usage before you compare headline rates, because the rate is the smallest part of the answer.
A product team budgets for video by taking the vendor’s per-minute rate and multiplying it by expected meeting hours. The number looks comfortable. They ship.
The first real invoice is roughly eight times the estimate. Nothing went wrong and the vendor did not overcharge. The team modelled sessions when the vendor bills participants, and every meeting had about eight people in it.
This is the most common costing mistake in video, and it is entirely avoidable. This guide explains how video call API pricing actually works, the four billing models you will encounter, the costs that never appear in the headline rate, and how to forecast your bill before you commit.
Table of Contents
ToggleHow is video call API pricing calculated?
Most video call APIs multiply participants by minutes by a rate. Ten people in a one hour meeting produce 600 participant minutes, not 60. Recording, transcription, and higher resolutions are usually billed separately on top of that base figure.
The industry standard formula is simple once you see it: participants multiplied by minutes multiplied by rate. Published analysis of video API pricing models puts the typical rate near 0.004 US dollars per video minute, though this varies by vendor and resolution.
The trap is that costs scale with group size, and group size is the variable product teams control least. As one analysis of usage-based pricing notes, many founders assume they pay per session when they actually pay per participant per minute. For group-based products, costs climb far faster than expected.
A worked example makes the scale obvious. These figures are illustrative, using the published 0.004 rate.
| Scenario | Sessions/month | Avg participants | Participant minutes |
|---|---|---|---|
| 1:1 support calls, 30 min | 500 | 2 | 30,000 |
| Team meetings, 60 min | 500 | 8 | 240,000 |
| Online classes, 90 min | 200 | 40 | 720,000 |
Same vendor, same rate, twenty four times the cost between the first and third row. Session count barely moved. Participant count did all the work.
The four pricing models you will encounter
Vendors rarely describe their model in these terms, so it is worth naming them yourself during evaluation.
| Model | You pay for | Predictable? | Best fit |
|---|---|---|---|
| Participant minutes | Every person, every minute | Low | Small calls, spiky usage |
| Monthly active users | Unique users per month | Medium | Heavy use by a stable base |
| Seats or plans | Licensed hosts | High | Internal use, known headcount |
| Self-hosted | Infrastructure and staff | High, but fixed | Sovereignty needs, large scale |
Participant-minute billing punishes exactly the thing many products want, which is large engaged sessions. Seat and plan pricing decouples cost from session size entirely, which is why platforms serving classrooms and all-hands meetings usually price that way. Convay is in that group, selling plans rather than metered participant minutes, so a session with 300 attendees does not cost thirty times a session with ten.
Neither model is better in the abstract. They are better for different usage shapes, and the only way to know which suits you is to model both against your own numbers.
What costs are missing from the headline rate?
Recording is usually billed per recorded minute and then again for storage each month it is retained. Higher resolutions often carry multipliers. Transcription, live streaming, and phone dial-in are typically separate line items. Compliance retention quietly turns storage into a permanent growing cost.
The base rate is the beginning of the bill, not the end. These are the additions that catch teams out.
- Recording, charged per recorded minute, then charged again for storage. Daily.co publishes a storage rate of 0.003 US dollars per minute alongside its recording rate
- Resolution and codec multipliers that raise the effective rate well above the advertised figure
- Transcription and AI summaries, usually priced separately from the call itself
- Retention policy, which is the real multiplier. Recordings kept for seven years never stop costing
- Phone dial-in and live streaming to social platforms, almost always extra
Retention deserves its own conversation with your compliance team before you sign anything. If your sector requires you to keep recordings for years, storage is not a small add-on. It becomes a permanent, compounding line item that grows every month you operate, and it is frequently larger than the call cost itself by year three.
How do you forecast your video API bill?
Estimate monthly sessions, average duration, and average participants, then multiply all three to get participant minutes. Add recorded minutes and multiply storage by your retention period. Model the same numbers at three times growth, because pricing models that work at launch often break at scale.
Build the forecast in five steps, and do it in a spreadsheet you can hand to finance.
- Count expected sessions per month.
- Multiply by average session length in minutes.
- Multiply by average participants. This is your participant-minute base.
- Add recording: what share of sessions get recorded, and what is the retention period in months.
- Repeat the whole calculation at three times your expected volume.
That last step matters most. Metered pricing is usually cheapest at low volume and most expensive at scale. Plan and seat pricing is the reverse. The crossover point decides which vendor is genuinely cheaper for you, and it rarely sits where the sales deck implies.
If you expect large sessions, check whether large-scale meetings are priced differently from standard ones. Many platforms treat them as a separate product with separate commercial terms.
When does self-hosting become cheaper?
Self-hosting removes per-minute costs entirely and replaces them with infrastructure and staff. The licence may be free. The operation is not.
Analysis of open-source video deployments makes this point directly: the cheapest tools by licence cost frequently become the most expensive once development and operations time is counted. Running production video means media servers, network traversal infrastructure, monitoring, scaling, and someone on call when a session fails during a board meeting.
Self-hosting usually wins in two situations. The first is very high volume, where per-minute charges exceed the fully loaded cost of running your own infrastructure. The second is regulatory, where data sovereignty requirements make external hosting unacceptable at any price. In the second case the cost comparison is not really the point, because the alternative is not permitted.
A middle path exists. On-premise deployment of a commercial platform gives you the sovereignty position without asking your team to become WebRTC infrastructure engineers.
Conclusion
Three things will keep your video budget honest.
Model participant minutes, never sessions. Group size drives your bill far more than meeting count does.
Price recording and retention as a separate long-term commitment, because storage you are required to keep never stops accruing.
Run the numbers at three times your expected volume. The cheapest vendor at launch is often the most expensive one the year after.
Frequently Asked Questions
Most vendors multiply participants by minutes by a rate. A one hour meeting with ten people produces 600 participant minutes, not 60. Recording, transcription, and higher resolutions are usually billed separately on top of that base figure.
A participant minute is one person spending one minute in a call. Four people in a fifteen minute meeting generate 60 participant minutes. It is the standard billing unit for metered video APIs, and confusing it with session minutes is the most common cause of underestimated video budgets.
The usual causes are modelling sessions instead of participant minutes, underestimating average group size, and overlooking recording storage that accrues every month a file is retained. Resolution multipliers and separately billed transcription also push the effective rate above the advertised one.
Usage-based pricing is generally cheaper at low volume and with small groups. Seat or plan pricing becomes cheaper as sessions get larger, because cost stops scaling with attendee count. Model both against your own forecast, then repeat at three times volume to find the crossover point.
Sometimes, but not as often as licence costs suggest. Self-hosting replaces per-minute charges with servers, bandwidth, and engineering time, and analysis of open-source deployments finds the cheapest licences often produce the highest total cost. It wins clearly at very high volume, or where data sovereignty rules make external hosting unacceptable.
Related guides
- Start with the complete video conferencing API guide
- Compare against the real cost of self-hosting
- See how architecture sets your cost ceiling
- Use a weighted scoring framework to compare vendors
Want a cost model built on your actual usage?
Bring your session count, average group size, and retention requirement. We will model it against plan pricing and show you where the crossover against metered billing sits for your volume.

