Production readiness
Validate credentials, settings, Webhooks, retries, monitoring, and contingency before taking your Asaas integration to production.
An integration validated in Sandbox is not automatically ready to operate in production.
Credentials, URLs, Webhooks, permissions, volume, monitoring, and operational behavior need to be reviewed in the production environment before go-live.
By the end of this page, you will have a checklist to validate whether the main components of the integration are truly ready to operate in production.
When to use
Use this checklist when:
- preparing a new integration for its first go-live;
- moving a flow validated in Sandbox to production;
- enabling a new feature in an existing integration;
- reviewing an integration before increasing its volume of operations;
- setting up a technical validation before releasing the flow to real users.
Before you start
Before production validation:
- complete functional testing in Sandbox;
- define which flows will be released at go-live;
- have access to the production environment's credentials and settings;
- identify who is responsible for monitoring the integration;
- define how the operation will be paused or recovered if a failure occurs after release.
How it works
Production readiness should validate each component again in the environment where the integration will actually operate.
%%{init: {"flowchart": {"nodeSpacing": 28,"rankSpacing": 32,"diagramPadding": 8,"padding": 10}}}%%
flowchart TD
A["Integration validated in Sandbox"] --> B["Review production environment"]
B --> C["Validate credentials and settings"]
C --> D["Test Webhooks and the full flow"]
D --> E["Validate failures and recovery"]
E --> F["Enable logs and monitoring"]
F --> G{"Go-live criteria met?"}
G --> GSim(("Yes"))
G --> GNao(("No"))
GSim --> H["Release integration"]
GNao --> I["Fix pending items"]
I --> B
H --> J["Monitor initial operation"]
classDef inicio fill:#DBEAFE,stroke:#2563EB,color:#1E3A8A,stroke-width:3px,font-size:17px
classDef processo fill:#E0F2FE,stroke:#0284C7,color:#0C4A6E,stroke-width:2px,font-size:17px
classDef decisao fill:#FEF3C7,stroke:#D97706,color:#78350F,stroke-width:2px,font-size:17px
classDef sucesso fill:#DCFCE7,stroke:#16A34A,color:#14532D,stroke-width:3px,font-size:17px
classDef recuperacao fill:#FEE2E2,stroke:#DC2626,color:#7F1D1D,stroke-width:2px,font-size:17px
classDef respostaSim fill:#22C55E,stroke:#15803D,color:#FFFFFF,stroke-width:3px
classDef respostaNao fill:#EF4444,stroke:#B91C1C,color:#FFFFFF,stroke-width:3px
class A inicio
class B,C,D,E,F,J processo
class G decisao
class H sucesso
class I recuperacao
class GSim respostaSim
class GNao respostaNao
linkStyle default stroke:#94A3B8,stroke-width:2px
The goal is not to repeat all the tests performed in Sandbox, but to confirm that the necessary dependencies and settings also exist and work correctly in production.
1. Review credentials and permissions
Before making any call in production, confirm that the application is using the credentials that correspond to the correct environment.
Check:
- whether the API Key in use was generated in production;
- whether any Sandbox credential remains configured;
- whether the credential is stored outside the source code;
- whether only the necessary services have access to the key;
- whether the available permissions match the operations executed by the integration;
- whether there is a defined procedure for replacing or rotating the credential.
Sandbox and production credentials belong to different environments. Always validate the origin of the API Key before go-live.
Avoid depending on undocumented manual changes at release time.
Sensitive settings must be separated by environment and be part of the application's deployment process.
2. Confirm environment URLs and settings
In addition to credentials, review all settings that change between Sandbox and production.
This includes, when applicable:
- the base URL used in requests;
- Webhook URLs;
- permissions;
- identifiers configured per environment;
- environment variables;
- integration-specific parameters;
- settings of external services related to the flow.
Do not assume that a setting created in Sandbox will also be available in production.
Each environment must be validated separately.
Keep Sandbox and production settings explicitly separated. This reduces the risk of a production application using credentials, URLs, or resources from the test environment.
3. Validate Webhooks in production
Webhook reception must be tested again in the production environment.
Confirm that:
- the correct URL is registered;
- the endpoint is externally reachable;
- security validations are working;
- the event is persisted before processing;
- the response occurs within the expected time;
- asynchronous processing is active;
- duplicate events can be handled without repeating effects;
- failures remain available for reprocessing.
Do not consider the flow validated just because the endpoint worked in Sandbox.
Infrastructure, network rules, domain, proxy, firewall, certificates, and intermediary services may be different in production.
4. Test the flow end to end
Run at least one complete flow using the production environment before releasing the integration for the full expected volume.
The test should include the relevant steps of the operation, such as:
- sending the initial request;
- storing the returned identifiers;
- tracking the entity's state;
- receiving the corresponding event;
- processing the update;
- updating the local record;
- confirming the final result of the operation.
%%{init: {"flowchart": {"nodeSpacing": 28,"rankSpacing": 32,"diagramPadding": 8,"padding": 10}}}%%
flowchart TD
A["Execute a controlled real operation"] --> B["Record ID and initial state"]
B --> C["Wait for update"]
C --> D["Receive and process event"]
D --> E["Update local system"]
E --> F{"Is the result consistent?"}
F --> FSim(("Yes"))
F --> FNao(("No"))
FSim --> G["Record validated flow"]
FNao --> H["Investigate mismatch"]
H --> I["Fix configuration or implementation"]
I --> A
classDef inicio fill:#DBEAFE,stroke:#2563EB,color:#1E3A8A,stroke-width:3px,font-size:17px
classDef processo fill:#E0F2FE,stroke:#0284C7,color:#0C4A6E,stroke-width:2px,font-size:17px
classDef decisao fill:#FEF3C7,stroke:#D97706,color:#78350F,stroke-width:2px,font-size:17px
classDef sucesso fill:#DCFCE7,stroke:#16A34A,color:#14532D,stroke-width:3px,font-size:17px
classDef recuperacao fill:#FEE2E2,stroke:#DC2626,color:#7F1D1D,stroke-width:2px,font-size:17px
classDef respostaSim fill:#22C55E,stroke:#15803D,color:#FFFFFF,stroke-width:3px
classDef respostaNao fill:#EF4444,stroke:#B91C1C,color:#FFFFFF,stroke-width:3px
class A inicio
class B,C,D,E processo
class F decisao
class G sucesso
class H,I recuperacao
class FSim respostaSim
class FNao respostaNao
linkStyle default stroke:#94A3B8,stroke-width:2px
Use controlled operations that are compatible with the flow that will be made available to users.
The goal is to verify the integration as a complete system, not just to confirm that an endpoint responds.
5. Also test failure scenarios
The success path is not enough to validate a production integration.
Include scenarios in which your application needs to deal with situations such as:
- validation error;
- temporary unavailability;
- timeout;
- inconclusive response;
- duplicate Webhook;
- failure while processing an event;
- delay in an update;
- need for a retry;
- mismatch between the local state and Asaas.
The goal is to confirm that the mechanisms defined in the previous chapters also work outside the ideal scenario.
A flow that works only when all calls and events occur as expected is not yet prepared to handle real integration failures.
6. Enable logs and traceability before go-live
Do not wait for the first incident to find out what information needs to be recorded.
Before releasing the integration, confirm that it is possible to correlate an operation across the different components of the system.
Record, when applicable:
- the internal identifier;
- the ID returned by Asaas;
externalReference;- date and time of the calls;
- the result of the requests;
- the known state of the entity;
- events received;
- the processing result;
- retries executed;
- failures and reconciliations performed.
Avoid logging credentials, sensitive data, or information that is unnecessary for diagnosis.
The goal is to be able to answer, from the available records, what happened to a given operation.
7. Configure monitoring and alerts
Logs let you investigate a problem. Monitoring lets you discover that it is happening.
Define indicators for the most important points of the integration, such as:
- an increase in errors in API calls;
- growth in
5xxresponses; - timeout occurrences;
- events with processing errors;
- events piling up in the queue;
- an increase in processing time;
- operations that remain in intermediate states longer than expected;
- recurring failures in retries;
- mismatches found by reconciliation.
Also define when a behavior should trigger an alert and who will be responsible for evaluating it.
Do not depend exclusively on a user complaint to discover a failure.
8. Review capacity and behavior under volume
An integration that works with few requests may behave differently when volume increases.
Before increasing traffic, evaluate:
- expected volume of operations;
- request concurrency;
- Webhook consumption speed;
- capacity of queues and workers;
- average processing time;
- retry strategy;
- limits applicable to the endpoints used.
The goal is to prevent an increase in volume from turning small failures into growing queues, simultaneous retries, or overload of the application itself.
9. Define a contingency plan
Before go-live, determine what will be done if a relevant failure occurs.
The plan should define, at a minimum:
- when to temporarily stop new operations;
- how to avoid new effects while the problem is investigated;
- how to identify affected operations;
- how to reprocess or reconcile pending operations;
- who should be contacted;
- what information should be gathered for investigation;
- which criteria allow the flow to resume.
The time to define this process is before the incident.
A contingency plan does not need to anticipate every possible failure. It needs to make clear how to reduce the impact, identify the affected operations, and safely recover the flow.
Go-live checklist
Before releasing the integration, validate:
| Item | Confirm |
|---|---|
| Environment | URLs and credentials correspond to production |
| Credentials | The API Key is protected and has the necessary permissions |
| Settings | Required parameters have been configured in the production environment |
| Webhooks | The production endpoint is active, tested, and monitored |
| States | The application tracks the necessary changes to entities |
| Events | Duplicates, delays, and failures can be handled |
| Retries | Timeouts and temporary errors do not create new operations without validation |
| Concurrency | The same operation cannot be executed simultaneously by multiple processes |
| Reconciliation | Mismatches can be identified and fixed |
| Logs | IDs, states, events, and attempts can be traced |
| Monitoring | Relevant failures generate indicators or alerts |
| Contingency | There is a procedure to pause, investigate, and recover the flow |
Attention — common mistakes
Avoid going to production:
- assuming the behavior observed in Sandbox will be identical without new validation;
- using an API Key from the wrong environment;
- keeping Sandbox URLs in production settings;
- without reviewing the permissions used;
- without registering or testing the production Webhooks;
- validating only the success path;
- without retry and duplicate prevention mechanisms;
- without enough logs to trace an operation;
- without failure monitoring;
- without defining who will monitor the first days of the integration;
- without a contingency plan.
Confirm your architecture
Before go-live, confirm that your team can answer:
- Do the credentials used belong to the production environment?
- Have the URLs and other settings been reviewed?
- Are the available permissions sufficient and appropriate for the flow?
- Have the production Webhooks been tested end to end?
- Can the application handle duplicate events and processing failures?
- Are timeouts and inconclusive responses handled without triggering unintended retries?
- Is there protection against concurrent executions?
- Has the flow also been validated in failure scenarios?
- Can mismatches be reconciled?
- Is it possible to trace an operation using the available logs?
- Are there alerts for the main types of failure?
- Does the team know what to do if a relevant failure occurs after go-live?
If any of these answers still depends on an investigation or decision during an incident, there is a pending production readiness item.
Best practices
- keep settings explicitly separated by environment;
- validate credentials, URLs, and Webhooks again in production;
- run a controlled end-to-end test before increasing volume;
- include failure scenarios in the validation;
- enable logs and monitoring before the first real operation;
- monitor the start of the operation more closely;
- increase volume gradually when the flow allows;
- keep owners and contingency procedures defined;
- periodically review the architecture as the integration evolves.
Being production-ready means confirming that states, events, Webhooks, retries, idempotency, reconciliation, and observability work in the real environment — and that the integration can recover when any of these steps fails.
Updated 2 days ago
