A2P SMS Operational Reporting: Metrics That Reveal Problems Without Hiding Uncertainty
A guide to defining metrics, denominators, and time periods; segmenting traffic carefully; and reporting pending or inconclusive DLRs without confusing them with failures or verified receipt.

What decisions an SMS operations report should help you make
An operations report is useful when it helps you decide what to review, compare, or escalate—not merely when it summarizes activity. It can help identify changes in volume, differences between segments, apparent delays, or inconsistencies in observed data.
Frame each metric as a question: What changed? In which population? Over what period? Based on what evidence? A change is a signal to investigate, not a causal explanation. For example, a drop in statuses reported as delivered does not, by itself, prove that a route failed; it may also be related to changes in traffic composition, reporting delays, or incomplete data.
- Identify the decision associated with each metric: investigate, compare, escalate, or continue monitoring.
- Include the data scope and observational limitations alongside the result.
- Avoid presenting a temporal correlation as a proven cause.

Define each metric before comparing: messages sent, acceptances, DLRs, and pending statuses
Before calculating rates, document what each metric counts and at what point in the flow it is recorded. “Sent,” “accepted,” “with a DLR,” and “delivered according to a DLR” are not interchangeable terms. The exact definition should match the events available in the platform and the applicable technical documentation; do not assume that all systems use the same semantics.
A report should distinguish attempts or messages sent from observed acceptances and delivery reports received. It should also distinguish conclusive statuses from pending, unknown, or contradictory ones. A received DLR is a status reported by the relevant system; by itself, it is not independent verification that the message appeared on the recipient’s device.
Maintain a shared, versioned glossary. If a definition changes, record when the change takes effect so that periods built using different rules are not compared.
- Specify the event that starts and ends each count.
- Clarify whether the data counts unique messages, attempts, or events; do not mix units.
- Document how retries, duplicates, and incomplete records are handled, based on the systems’ actual capabilities.
- Identify statuses reported by third parties and avoid calling them verified receipt.

Choose comparable denominators and time windows
Every rate needs an explicit denominator. For example, a proportion of reported delivery statuses should specify which messages it includes and which it leaves out: those sent during the period, those accepted, or those that have already received an update. These bases answer different questions and can produce different figures.
Also define how time is assigned: by the time of sending, acceptance, or receipt of the report. When comparing periods, keep the same time zone, duration, and inclusion rule, or explain the difference. For recent messages, some delivery reports may not have arrived yet; comparing a mature cohort with one whose tracking is still open can distort the result.
The available evidence does not establish a single window that is right for every case. Set an observation window that fits your reporting times and processes, and publish it as internal methodology, not as a universal standard.
- Write the formula and denominator alongside the rate’s name.
- Record the period, time zone, and event used to assign each record.
- Separate cohorts with complete follow-up from those that may still receive updates.
- If you change a window or inclusion rule, note the change in the report.
Segment by destination, operator, sender, traffic type, and route without creating misleadingly small groups
Segmentation can help identify where a change is concentrated. Depending on which fields are actually available and reliable, analyze by destination, operator, sender, traffic type—for example, OTP, transactional, or legitimate marketing—and route. Do not assume that every system has all these fields or that their values are normalized.
Start with an overall view and add dimensions gradually. If several factors change at once, an observed difference does not show which one explains it. Where possible, compare similar periods and record known operational changes, such as configuration changes or shifts in traffic mix.
Small groups can produce highly volatile percentages and encourage premature conclusions. The available evidence does not establish a universal minimum group size. Set internal reporting rules based on volume, stability, and privacy requirements; when a group is too small to interpret, combine it cautiously or mark the data as insufficient.
- Check the quality and consistency of each field before using it for segmentation.
- Show volume alongside the percentage to reveal the size of the base.
- Avoid comparing groups with very different compositions or periods without noting the difference.
- Do not infer causality from a coincidence between route and outcome.
Show latency with percentiles and distributions, not just averages
An average compresses all values into a single number and can hide the fact that some messages take much longer than others. When analyzing timing, consider showing the distribution and percentiles alongside the average, provided the data and volume support reliable calculations.
Define precisely which interval you call latency—for example, the time between two events recorded by your systems. Do not combine measures with different start and end points, or present the time to a DLR as if it necessarily measured the time until receipt on the device. A delivery report may arrive late or may not be available, which affects what can be concluded.
Accompany percentiles with the number of observations, the period, and the population analyzed. If there are few records or missing values, state that limitation instead of giving a false impression of precision.
- Document the events that define the start and end of each timing measure.
- Show distributions and percentiles alongside volume and missing data.
- Do not equate reporting latency with latency until verified receipt.
Separate connectivity availability, delivery outcomes, and the quality of observed data
Connection availability, message acceptance, reported delivery statuses, and data completeness describe different aspects. A system may be available while delivery results still warrant investigation; messages may also have no conclusive status even when observed connectivity has not changed.
Organize the report into separate sections and specify the source of each data point. Do not infer delivery quality from a connectivity signal alone, or connectivity quality from DLRs. If observability depends on external systems, explain which events may be missing or delayed.
Keeping these areas separate helps determine the next step: review connectivity, investigate reported outcomes, or validate record integrity. It does not, by itself, prove the cause.
- Label connectivity, acceptance, delivery outcomes, and data quality separately.
- State which system records each event and what coverage limitations apply.
- Treat any relationship between indicators as a hypothesis to validate.
Represent unknown, late, and contradictory statuses without turning them into certainties
A pending or inconclusive DLR should not automatically be turned into either a failure or a confirmed delivery. Show it as a visible category and define it clearly for the reader, based on the status your systems actually report. If a report can be updated later, preserve the distinction between the current observation and the final outcome, if and when that outcome arrives.
If signals conflict for the same message, do not silently choose the more favorable or unfavorable one. Define a documented reconciliation rule if the system allows one; otherwise, retain the case as contradictory or unresolved and exclude it from calculations that require a conclusive classification, explaining the effect of that exclusion.
The available evidence does not establish a universal taxonomy or technical rules for resolving every DLR status. Consult the applicable platform specifications and documentation before interpreting specific codes.
- Explicitly separate confirmed according to the report, pending, unknown, and contradictory statuses, if these exist in your data.
- Show how many records fall into each category and how they affect rates.
- Do not call a pending status “failed” or a DLR “verified receipt.”
- Record reconciliation rules and status changes to maintain traceability.
Add volume, operational changes, and limitations to every KPI
A key performance indicator (KPI) in isolation can be misleading. Include the base volume, period, denominator, proportion of inconclusive statuses, and any known operational changes that affect comparability. If traffic combines different use cases or segments, describe that composition before assigning significance to a change.
Turn findings into investigative questions. If a rate changes for a specific destination and period, first verify the definitions, volume, and maturity of the data; then compare segments and consult relevant operational records. Keep conclusions at the level supported by the evidence: a signal, an observed pattern, or a confirmed cause.
As industry context, the A2P market report prepared for Telefónica by Analysys Mason describes differences between A2P and P2P traffic profiles and interconnection challenges. This is a reminder that composition and interconnection matter when interpreting metrics, but it does not allow a specific change to be attributed to a route, operator, or event without specific evidence.
- Minimum KPI template: definition, numerator, denominator, period, volume, and excluded statuses.
- Add the proportion of pending or inconclusive records and the data source.
- Note known operational changes without presenting them as proven causes.
- End each finding with a question and the next validation step.
Frequently asked questions
Does a DLR confirm that the SMS reached the phone?
A DLR is a status reported by the relevant system. By itself, it is not independent verification that the message appeared on the recipient’s device. The report should describe what it actually observes and avoid claiming more than the data supports.
How should a pending DLR be counted?
Show it as pending or inconclusive, depending on the terminology that matches your data. Do not automatically count it as a confirmed delivery or a failure. State how many cases there are, which observation window you use, and how this category affects calculations.
Which denominator should be used for a delivery rate?
It depends on the question and the events available. Define whether the base is messages sent, accepted, or those that already have a reported status. Publish the formula and avoid comparing rates with different denominators as if they were equivalent.
What is the minimum size for a segment to be reported?
The available evidence does not support a universal threshold. Set an internal rule suited to volume, statistical stability, and privacy requirements. Show the size of the base and mark groups as insufficient when they do not allow cautious interpretation.
Does a deterioration in one segment prove that the route is the cause?
No. It is a signal to investigate. First check whether periods are comparable, whether metric definitions are consistent, the traffic composition, the maturity of statuses, and any known operational changes. Attribute causality only when there is sufficient specific evidence.
Sources consulted
- 3GPP specifications3GPP
- ITU-T E.164International Telecommunication Union
- GSMA resourcesGSMA
- SMS A2P - Telefónica Global SolutionsTelefónica Global Solutions
- Informe para Telefónica: El mercado de mensajería A2PTelefónica / Analysys Mason