The branch called at eleven forty: the video call with the client had dropped three times that morning. You opened a ticket. Two days later it came back closed with the usual sentence, that the circuit remained within contracted parameters. And there on your screen is your SD-WAN graph, with three loss spikes sitting right in that window. Nobody is lying. You and your carrier are measuring different things, and the differences are nobody's interpretation: they are written into the standards underneath your contract.
Between which two points
The performance metrics of an IP network are not properties of "the network": they are defined between two measurement points. One-way delay (RFC 7679) and one-way loss (RFC 7680) are defined from a specific source to a specific destination. Change the two points and the number changes, without anyone having falsified anything. The same holds at the Ethernet layer with Y.1731: the carrier sends measurement frames between two MEPs that it configures — DMM for delay, LMM or SLM for loss — and the result describes the stretch between those two ports.
The carrier's points are the edges of its network: usually its access device and the handoff interface at your site. Whatever falls outside that pair — your LAN, your firewall, the internet transit beyond the handoff, the SaaS at the far end — is not in the number and never claimed to be. Your SD-WAN measures something else, equally legitimate: the whole tunnel, from the site appliance to the egress node, including stretches the carrier never sold you and cannot guarantee.
The question "was the circuit fine?" has no single answer until you say between which two points. Until you do, both sides can be right at the same time, and they will be.
Over how long
ITU-T Recommendation Y.1541 is the one that sets quality objectives for IP-based services. Its class 0 — the demanding one, the one that covers real-time voice and video — calls for a mean transfer delay (IPTD) below 100 ms, a delay variation (IPDV) below 50 ms and a loss ratio (IPLR) below 1·10⁻³, that is 0.1%. It suggests a one-minute evaluation interval, requires that the interval be recorded alongside the observed value, and adds a sentence worth reading slowly: any minute observed should meet those objectives. Any of them. Not the average of the month's.
IPDV, moreover, is defined as the upper bound on the 1−10⁻³ quantile of IPTD minus the minimum IPTD. In plain terms: one packet in a thousand may fall outside and the objective is still met, by definition. It is the correct way to describe a network statistically, and it is still enough for an engineer to be looking at a compliant report while you were hearing the call break up.
Now look at your contract. In every access contract we have been handed to review, what gets committed is a monthly availability percentage. The per-minute class is left out. They are two different products and the second is far cheaper to meet, because a month-long average absorbs episodes that a single minute does not. The arithmetic fits in a table. A thirty-day month has 2,592,000 seconds:
| Contracted availability | Allowed downtime per month | 20-second drops that fit |
|---|---|---|
| 99,5 % | 3 h 36 min | 648 |
| 99,9 % | 43 min 12 s | 129 |
| 99,95 % | 21 min 36 s | 64 |
| 99,99 % | 4 min 19 s | 12 |
Look at the 99.9% row, the most common one in business access contracts. A twenty-second drop consumes 0.00077% of the month. Inside that 99.9%, 129 of those drops fit: more than four a day, every day of the month. If they land during office hours and on top of calls, the branch has had a hellish month, the monthly report comes out green, and both of those are true at once. The percentage is measuring something else.
Two more filters usually sit on top of that, and we say this as a reading of the contracts people send us, not as an industry statistic: a minimum duration below which an event does not count, and a clock that starts when you open the ticket rather than when the problem starts. Go to your contract and look for those two clauses. They are almost always there, and between them they swallow most of the episodes that genuinely annoy people.
Which queue the probe travels in
For continuous measurement, many carriers use TWAMP, the two-way active measurement protocol described in RFC 5357. In TWAMP a test session carries a DSCP, and the RFC uses a MUST: the reflector has to use that same DSCP in the packets it sends back. The measurement is, by design, per traffic class. And that is a good thing, because it is exactly what you need in order to verify a contracted class of service.
The problem starts when the probe's class and your traffic's class are not the same. At the handoff, DSCP remarking is routine: what left your appliance marked EF can enter the carrier network as best effort if the mapping was never agreed in writing. The probe then reports, entirely truthfully, an empty queue, while your video waits in a different one that is full. Nobody cheated; a per-class measurement is simply being read as a statement about everything on the wire.
Then there is sampling. An active probe does not watch continuously: it sends bursts of packets at a configured cadence, and between bursts it is watching nothing. Our own SmokePing install, to give the number we can actually give, runs by default at twenty pings every five minutes. At that cadence, a queue that fills for eight hundred milliseconds may not coincide with any sample at all. Nobody was fooled: nobody was looking at that instant. Before citing a probe in a claim, find out how often it measures.
The turn-up test certifies turn-up day
When the circuit was turned up they sent you a PDF with graphs and a "pass" stamp. Most likely it was a Y.1564 test, the ITU-T Ethernet service activation methodology. It has two phases. The first verifies configuration by stepping traffic to 25, 50, 75 and 100% of the committed rate, then to the excess rate and then above it, checking that policing behaves the way the contract says. The second holds 100% of the committed rate and measures loss, delay, variation and availability. The three test periods the Recommendation requires support for are fifteen minutes, two hours and twenty-four hours: fifteen minutes for a metro service over a network already carrying traffic, two hours for single-operator long haul, twenty-four for an international service crossing several networks.
It is a good test and it does exactly what it promises: certify the service at activation. But look at the scale. The most common case, the metro one, gets certified on fifteen minutes of traffic, and you will be living with that circuit for five years. The PDF tells the truth about that quarter of an hour. It says nothing about last Tuesday at eleven forty, and never intended to. Producing it in a dispute — in either direction, because we produce it too when it suits us — is asking a document for something it does not contain.
Your SD-WAN does not hold the whole truth either
We are writing this section as the people who deploy SD-WAN, and we are writing it anyway. An argument that only stands up when the box vendor makes it is worth very little. Your orchestrator's telemetry is also a probe, with its own cadence, packet size and class, and it carries three biases of its own that are worth knowing before you take it into a meeting.
- It measures the whole path, not the circuit. That includes tunnel encapsulation and an internet middle that nobody sold you with any commitment. A bad path does not imply a bad access circuit, and if you present it as though it did, the carrier is right to hand it back.
- When it acts, it changes what it was measuring. As soon as the appliance steers traffic off the bad path, that path stops carrying load and stops looking bad. The "after" graph proves the steering worked, not that the circuit healed. Those are two different conclusions and it is very easy to claim the second without having earned it.
- The evidence expires before the claim does. Check how long your orchestrator keeps the series at second resolution before compacting it into averages. If your claim window is thirty days and the detail dies at seven, your evidence evaporated before the argument did. Nobody will warn you: it is in the retention settings and you have to go and look.
None of this takes anything away from a well-deployed SD-WAN; it says what it is for and what it is not. We wrote back in July about when it pays off and when it does not, and that still stands: it solves a specific problem, and without that problem it is one more layer to maintain. What we are adding today is that its telemetry, one of its best reasons to buy, needs context before it becomes evidence.
Seven lines of annex before you sign
Seven sentences to write into the technical annex of the contract, or to ask by email before signing and keep the answer. Each one closes one of the ambiguities above. If the account manager cannot answer six out of seven, you already know something useful.
- Between which two points the measurement runs, named by the specific interface rather than "the customer network". If the point on your side is the carrier's device at your site, write it down that way.
- Which parameter and against which standard: delay, variation and loss with exact values, citing Y.1541 or Y.1731 where applicable. A "guaranteed quality" with no parameter is decorative.
- The evaluation interval, and whether the figure is a mean or a percentile — and in the latter case, which one. The difference between a five-minute mean and a one-minute 95th percentile is the difference between seeing the episode and missing it.
- Which class the probe travels in, and what marking your traffic receives at the handoff. If you are going to send video marked EF, get it in writing what the carrier does with that marking at its boundary.
- What counts as an event: minimum duration, and whether the clock starts at the carrier's own detection or when you open the ticket. Of everything on this list, it is the line we have most often seen decide how the argument ends.
- Who provides evidence, in what format, and how long each side retains it. If they keep ninety days and you keep seven, in practice only one party provides evidence.
- What the remedy is and what it does not cover. In every contract we have reviewed it has been a percentage of the monthly fee. It is good that it exists, and worth saying out loud what you already know: it does not pay for the lost meeting. If the service is critical, the money is better spent on a second path than on negotiating the penalty.
From both ends of the phone call
We sit in an odd position that happens to help here: we are a network operator in our own right — BGP, transit, peering and a public looking glass at lg.everywan.com — and at the same time we run managed services for companies whose circuits somebody else sold them. We have written the claim and we have written the reply. Bias up front: it suits us if you buy managed services, so assume it. What follows works exactly the same if you build it yourself.
Two habits come out of all this and they have saved us arguments. First: keep the spread, not just the mean. We use SmokePing precisely because it preserves the distribution of each probe round, so a twenty-second episode leaves a visible mark that a monthly average erases without trace; and Zabbix for thresholds and for the alert that genuinely wakes somebody up, which is a craft of its own and one we have written about. Second: give the circuit a name. We keep NetBox as our source of truth with its reference, its handoff interface and its owner, because a good share of the claims that die, die in the first email, arguing about which circuit we are even discussing.
And a third, less comfortable one: some outages no annex can fix, because what decides the outcome is where you are able to act. We covered this when we wrote about why a local blackhole will not save your access link: if the link is already full by the time traffic reaches your door, whatever you do at your door arrives late. Knowing who owns each stretch is useful for two things, and claiming is the less important one; the other is knowing who to ask to act.
What you are actually buying
An availability SLA works like insurance: it hands back part of the fee when the service fails for a long stretch. It has value, it is good that it exists and there is nothing murky about it. What it does not do is describe how your network behaves on Tuesday at eleven forty, because that would mean measuring something else, somewhere else, over a different window — and that is precisely what was not contracted.
We ran into the same mechanism in August with a Microsoft 365 incident: search had been broken for days and, read against the SLA's own Downtime definitions, it did not count as an outage. There it was the service definition doing the work; here it is the measurement point and the window. The substance is identical: what a contract treats as a failure is decided by the contract, not by the user whose call just dropped.
Our usual line says that failure is inevitable and an outage is a design decision. The argument with the carrier, on the other hand, is not inevitable at all: it is the consequence of nobody writing down, before signing, what "this works" is supposed to mean. Seven lines, half a page, written once.
If you are building or renegotiating a multi-site WAN, that annex is worth more than a price comparison: it decides who is right eight months from now. It is where we start when we design an SD-WAN or take over a company's networks and communications, before touching a single box. If you want us to look over your current contract and tell you which of the seven lines is missing, get in touch.
Sources. Class 0 objectives (IPTD 100 ms, IPDV 50 ms, IPLR 1·10⁻³), the definition of IPDV as the upper bound on the 1−10⁻³ quantile of IPTD minus the minimum IPTD, the one-minute evaluation interval and the requirement that any minute observed should meet the objectives: ITU-T Recommendation Y.1541, "Network performance objectives for IP-based services", Table 1 and clause 5.3.2. The same class 0 is summarised in RFC 5976 §2.1. Two-phase activation methodology, the configuration-test steps and the three mandatory test periods of fifteen minutes, two hours and twenty-four hours with their selection criteria: ITU-T Recommendation Y.1564, "Ethernet service activation test methodology", clause 8.2. Measurement frames between MEPs (DMM for delay, LMM and SLM for loss): ITU-T Recommendation Y.1731. DSCP marking on test sessions and the reflector's "MUST" to reuse it: RFC 5357 (TWAMP). One-way delay and loss metrics defined between a source and a destination: RFC 7679 and RFC 7680. The availability arithmetic is ours, over a thirty-day month (2,592,000 s), and can be redone with a calculator. What we say about how contracts are written — the minimum duration, when the clock starts, the remedy as a percentage of the fee — is our reading of the contracts we have been handed to review, not an industry statistic, and it is flagged as such in the body. The capabilities we cite — our own network with BGP, transit and peering, a public looking glass at lg.everywan.com, SmokePing, Zabbix, NetBox, multi-site SD-WAN and managed services — are in our public documentation; what we say about having been on both sides of a claim is our own account of our work and includes no customer data. Cover photo: "North Central Telephone Cooperative Rural Broadband keeps TN and KY High-Speed - RUS (20180927-RD-LSC-1040)", USDA, public domain, via Wikimedia Commons, cropped.
Do you know between which two points you are measured?
When we design a multi-site SD-WAN we start with the measurement annex and the circuit inventory, not with the first appliance. If it turns out that what you have too much of is SD-WAN and what you are missing is a second access circuit, we will tell you so anyway.
Talk to everyWAN