Back to blog SMS Quality

Synthetic Tests and Real A2P SMS Traffic: How to Combine Evidence

Synthetic tests check controlled flows; production shows how real traffic behaves. Learn how to combine both sources of evidence and interpret their limitations.

Diagram comparing synthetic tests with observations of A2P SMS production traffic

Two Sources of Evidence, Two Different Questions

A synthetic test generates test traffic under deliberate conditions to observe a route or verify a technical flow. A sandbox can simulate traffic in an isolated environment, making it possible to check features, integration, and edge cases. This activity alone does not prove that an SMS reaches a real subscriber.

Production monitoring analyzes messages sent in normal operation. It can show how real traffic behaved over a given period and under specific conditions, but it does not automatically explain why a result occurred. Neither method on its own demonstrates how all recipients, operators, senders, or traffic types will behave.

  • Use synthetic tests to check whether a controlled flow works and to detect changes under repeatable conditions.
  • Use production data to understand observed behavior for real messages in the relevant operational context.
  • Interpret both sets of results in light of their sample, configuration, and time period.
Two Sources of Evidence, Two Different Questions

Clarify What You Want to Find Out Before Testing

The question determines which method may be useful. A controlled test can help check whether an HTTP or SMPP integration sends a request in the expected format. To assess observed availability, DLR consistency, latency, or sender and content behavior, first define which signals will be observed and in what context.

It is useful to distinguish the availability of a technical flow from the delivery outcome. It is also important to distinguish a DLR reported by the route from a signal that independently verifies receipt on a device. A reported acknowledgment does not, by itself, confirm receipt on the handset.

  • Formulate a specific question, such as whether a valid request triggers the expected flow or whether an observed metric has changed.
  • Define what result you will observe, which system records it, and what uncertainty remains.
  • Do not present simulated statuses or reported DLRs as independent proof that a message was received on a phone.
Clarify What You Want to Find Out Before Testing

Design a Controlled and Legitimate Synthetic Test

Before sending messages, define the objective and keep constant any variables you do not intend to evaluate. Record the destination, operator where it can be determined, sender, traffic type, legitimate content, time, route, and configuration. Changing several dimensions at once can make a difference harder to interpret.

Use only authorized test numbers and legitimate content. A sandbox with verified numbers lets you test flows and edge cases, but it should not be confused with sending messages to real subscribers. If the evaluation requires observing a real network, agree on an authorized procedure in advance and follow applicable rules.

  • Define a hypothesis and comparison criterion before running the test.
  • Where possible, change one variable at a time and preserve the configuration so you can repeat the evaluation.
  • Record the destination, known operator, sender, message type, content, time, and available route or environment identifier.
  • Distinguish sandbox results, controlled tests on real networks, and normal production traffic.

Compare Compatible Metrics, Not Similar-Sounding Labels

A useful comparison requires signals to have the same meaning. Record accepted requests, reported delivery statuses, observed latency, and flow availability separately. If there is an independent signal confirming receipt, document how it was obtained; do not combine it with reported DLRs.

Align the time periods, the definition of each metric, and the sending conditions. End-to-end latency may not be comparable with the time until a DLR is received. Likewise, a simulated result should not be combined with results from real subscribers as though they came from the same population.

  • Specify the start and end points of each latency measurement.
  • Keep reported DLRs separate from verifiable receipt signals.
  • Compare observations only when their definitions, environments, and time windows are compatible.
  • Identify missing or incomplete data; do not automatically classify it as a successful or failed delivery.

Interpret the Synthetic Sample with Care

A test involving a small number of numbers or runs describes what was observed in that sample, not necessarily how an entire population behaves. Test numbers may not represent the diversity of subscribers, devices, plans, operators, or usage conditions. A sandbox test is more specific still: it checks what that environment allows you to simulate.

Repeating tests can help identify whether an observation is reproducible, but it does not by itself eliminate selection bias or prove that the result applies to other destinations or traffic types. Describe the sample size and composition, the time period, and known limitations.

  • Avoid extrapolating a local result to all users or an operator’s entire coverage area.
  • Do not interpret repeatability as a universal delivery guarantee.
  • Explain which people or conditions are outside the sample.

Segment Production Data and Protect Personal Information

Aggregated metrics can conceal important differences. When the available data allows, analyze by destination, operator, sender, and traffic type, as well as across comparable time windows. Segmentation can help identify where a variation appears, but does not by itself prove its cause.

Limit analysis to the data needed to answer the operational question. Consider access restrictions and data minimization or aggregation in line with applicable obligations. Avoid including phone numbers, message content, or personal identifiers in reports that do not need them, and define how test data will be handled.

  • Keep destination, operator, sender, and traffic-type segments separate when relevant.
  • Consider using identifiers or aggregation that reduce exposure of personal data.
  • Limit data access and retention according to the purpose and applicable rules.
  • Do not use an HLR lookup as proof of consent, identity, ownership, or guaranteed delivery.

Investigate Differences Without Jumping to Conclusions

If a synthetic test and production data disagree, first check whether they measure the same thing. Verify the environment, time period, destination, operator, sender, content, volume, and definition of the result. Also check for incomplete data, configuration changes, or differences between simulation and real traffic.

A discrepancy is a signal to investigate, not automatic proof that a particular route, operator, or change caused the result. Consider alternative explanations, review the available evidence, and repeat a controlled test if it could provide useful information. Record what is known, what is not, and what action was taken.

  • First confirm that the metrics and time windows are comparable.
  • Check whether any route, destination, sender, content, or configuration variable changed.
  • Review patterns in related segments before attributing a cause.
  • Repeat the evaluation under documented conditions and communicate the uncertainty.

A Practical Framework for Choosing Which Evidence to Use

To validate an integration or reproduce an edge case, start with a synthetic test or sandbox and make clear which aspects are simulated. Sinch’s sandbox documentation describes an isolated environment where traffic is simulated using verified numbers and notes that its endpoints and webhooks match those in production. This can help test the integration, but it does not prove that production conditions or delivery outcomes are identical.

To assess the behavior of real messages, analyze segmented production data and explain what signal each metric represents. If a decision affects a route or operation, consider both perspectives when relevant without treating one as a substitute for the other. Document the question, method, conditions, sample, metrics, limitations, and criteria for repeating or closing the evaluation.

  • Synthetic testing: can help with integration, reproducibility, and controlled scenarios.
  • Production: allows you to observe real traffic behavior in the segments and periods analyzed.
  • Combined evidence: can inform different parts of a decision if the limits of each source are documented.
  • Repeat or extend the analysis if the sample does not answer the question, the metrics are not comparable, or the discrepancy remains unexplained.
FAQ

Frequently asked questions

Does a synthetic test confirm that an SMS will reach all recipients?

No. It describes what was observed under the tested conditions and within the tested sample. A sandbox can simulate traffic and help test flows, but it is not, by itself, evidence of delivery to real subscribers or proof of universal results.

Is a DLR equivalent to verified receipt on a phone?

Not necessarily. A DLR is a status reported by the system or route. It should be distinguished from an independently verified receipt signal, and the uncertainty of each signal should be documented.

What should be segmented in production metrics?

Where the data allows and it is relevant, analyze by destination, operator, sender, and traffic type, using comparable time windows as well. Segmentation can reveal patterns, but does not establish the cause on its own.

When should an evaluation be repeated?

Repeat it if the sample does not answer the question, a relevant condition has changed, the metrics are not comparable, or a discrepancy persists. Record the conditions so the new result can be interpreted.

Sources consulted

  1. Especificaciones 3GPP: series de especificaciones3GPP
  2. Recomendación ITU-T E.164International Telecommunication Union
  3. Recursos de redes móvilesGSMA
  4. Sandbox de SMS: pruebas de API y simulaciónSinch
  5. Guía del usuario de AWS End User Messaging SMSAmazon Web Services