Partial Degradation at an A2P SMS Provider: How to Isolate the Impact and Choose an Operational Response
A cautious guide to identifying whether an SMS issue affects only certain segments, gathering comparable evidence, and deciding whether to monitor, limit, or suspend traffic.

What Partial Degradation Means
Partial degradation is an operational hypothesis: some parts of the traffic show different results from others, while aggregate data may hide that difference. The pattern could be associated with a destination, operator, sender, message type, or time window, but the available data alone cannot establish the cause.
Do not assume that the provider is completely down or that the issue originates in its network. A change may also be related to the source system, traffic configuration, or observability. The first goal is to define what changed and which set of messages was affected.
- Treat the incident as a hypothesis to investigate, not a confirmed diagnosis.
- Compare equivalent segments and comparable time periods before changing traffic.
- Separate what you observed from the cause you attribute and the action you propose.

Segment Before You Decide
Build a view that lets you compare affected traffic with traffic that appears normal. Use dimensions present in your records and relevant to operations. Avoid combining different categories into a single percentage: a global average can conceal a concentrated issue.
Keep the conditions of the comparison consistent where possible. If the destination, sender, and traffic type all change at once, it will be harder to determine which dimension coincides with the degradation. Segmentation helps define the pattern, but does not, by itself, prove its cause.
- Destination and operator, when this information is available and reliable.
- Sender and traffic type, such as OTP, transactional messages, or consented campaigns.
- Provider or route, observed status, and time window.
- Volume and denominator for each segment, so isolated counts are not compared as if they were equivalent rates.
- Recent configuration or traffic changes that may affect interpretation.

Compare Acceptance, Errors, DLRs, and Independent Signals
Gather the available signals and define what each one represents. A request accepted by an interface is not the same as delivery to a handset. Likewise, a received DLR is a status report, not independent confirmation that a person received or read the message on their device.
Compare send records, errors, and delivery reports with any genuinely independent observations you have, such as controlled test results or measurements from your own systems. Do not claim that one signal validates another unless they measure the same event.
The available evidence does not specify how SMS statuses are technically signaled or interpreted across all networks. Therefore, document the source and operational meaning of each field according to the applicable interface and agreements, rather than inferring protocol details.
- Record submitted and accepted requests, observed errors, and received DLRs by segment and time period.
- Note the definition each system uses for every status; do not assume two providers label statuses identically.
- Distinguish delivery reports from independent checks of handset receipt.
- Flag missing, delayed, or non-comparable data; the absence of a DLR does not, by itself, prove a cause.
Assess Impact Based on Message Use
Not all messages have the same urgency or operational consequences. Separate OTPs, transactional messages, and consented campaigns when assessing the impact on the service and its recipients. Do not extrapolate the behavior of one segment to another.
Agree with service owners on what delay, error, or uncertainty requires intervention. The threshold depends on the message’s function and your organization’s policies; the available evidence provides no universal value that applies to every case.
- OTP: assess how a delay or failure affects the flow that depends on the code.
- Transactional messages: identify which process or communication is affected and for how long.
- Campaigns: check the impact on consented sending and apply the relevant rules and controls.
- Prioritize assessment by impact and scope, not just total volume.
What to Ask the Provider
Share a concise, reproducible summary of the observed pattern. Ask the provider to confirm which segments and time periods it was able to review, what status it assigns to the incident, and which known constraints might be relevant. Avoid asking for a definitive cause before there is supporting evidence.
Agree on the next update and the data format that will allow results to be compared. If privacy, retention, or access limits apply, use the channels and fields authorized under your operational relationship.
- Scope investigated: destinations, operators, senders, traffic types, and time window.
- Incident status and the observations supporting it.
- Known constraints or relevant changes the provider can confirm.
- Next steps, who is responsible for the update, and the agreed time to review it.
- Definitions of reported statuses and any limitations of the shared data.
Choose Whether to Monitor, Limit, or Suspend
The response should be proportionate to the impact, strength of the evidence, and reversibility of the action. If the pattern is uncertain and the impact appears limited, enhanced monitoring may be reasonable under your internal controls. If the issue is well contained within a segment and a reversible action is available, consider limiting only that traffic. If the impact is high or you cannot control the risk with a narrowly scoped measure, consider suspending the affected segment or route in accordance with your procedures.
There is no universal numerical threshold or technical rule here that automatically determines when to act. Define who has authority to make each decision and document the conditions that would justify changing it.
- Monitor: the pattern is not yet confirmed and the risk can be monitored safely.
- Limit: the impact appears localized and the action can be scoped and reversed.
- Suspend: the impact or uncertainty exceeds what your controls allow you to accept.
- In each case, assign an owner, the next review, and an escalation condition.
Evaluate an Alternative as Mitigation, Not a Cure
An alternative route could reduce exposure to the affected segment, but it does not establish the cause or guarantee that the issue will not occur there as well. The available evidence does not support claiming that an alternative route will resolve a degradation.
Before moving traffic, confirm that the alternative is authorized for the use case and that you can monitor its results using comparable criteria. Apply the relevant consent and compliance controls. Document the change as temporary mitigation until there is enough data to review the decision.
- Define which traffic would be moved and who authorizes the change.
- Set evaluation signals that are comparable with those for the original segment.
- Agree on a review condition and a way to reverse the measure.
- Do not present an observed improvement as proof that the underlying cause has been resolved.
Record the Decision and Conditions for Resuming Traffic
Keep a concise record that explains what was observed, what was decided, and why. Include the segments and time periods reviewed, the sources consulted, uncertainties, and responsible people. This helps prevent temporary mitigation from being mistaken for a confirmed resolution.
Define in advance what evidence will allow traffic to resume and who will verify that the conditions are met. Resumption should follow applicable internal controls and operational agreements; do not base it solely on a DLR or a claim whose scope has not been defined.
- Incident: scope, time periods, available signals, and missing data.
- Decision: maintain, limit, or suspend, with the rationale and owner.
- Communications: provider, internal teams, and the next agreed update.
- Review: conditions for maintaining or reversing mitigation and criteria for resumption.
- Closure: record whether the cause was confirmed, remains unknown, or the impact was only mitigated.
Frequently asked questions
Does a DLR confirm that a message reached the phone?
It should not be treated as independent confirmation of receipt on the handset. It is a status report whose interpretation depends on its source and context. Always distinguish a DLR from an independent check of receipt.
How can I tell whether a failure is partial or general?
Compare segments defined by destination, operator, sender, traffic type, and time period, based on the available data. If some segments behave differently from others, the pattern may be partial, but segmentation alone does not identify the cause.
Should I switch to another route immediately?
Not necessarily. First assess impact, evidence, and reversibility. An alternative may serve as mitigation, but it does not guarantee delivery or prove that the original cause has been resolved.
What should I share when opening an incident with the provider?
Send the scope and time periods investigated, the signals observed, relevant status definitions, and data limitations. Ask the provider to confirm what it was able to review, any known constraints, and the next steps.
Is there a universal threshold for suspending traffic?
The available information does not support a universal numerical threshold. Define internal criteria based on the use case’s impact, confidence in the evidence, ability to limit the scope, and reversibility of the action.
Sources consulted
- 3GPP specifications3GPP
- ITU-T E.164International Telecommunication Union
- GSMA resourcesGSMA