How you learn that something broke
We have no auto-updating public status page. Instead there is a clear order of things: who calls whom, what we count as an outage, and what happens once it is fixed.
A green dot on a status page is worth nothing when the nightly export has stalled and the page does not know it. We prefer a message to the network owner — from a person who is already digging in.
How you learn about a failure
A failure affecting the calculation or order sending reaches the owner from us — in the same messenger the orders travel through. We do not wait for the shelf to notice.
If data is older than it should be, that shows next to the numbers, not in a separate system-health section. Decisions are made on the numbers, so the warning belongs beside them.
An update that could touch the nightly calculation or a send is agreed beforehand, not announced afterwards. A network's ordering schedule is stricter than our wish to ship.
What counts as an incident
Not only a service outage. An incident is anything that makes a decision run on a wrong picture, or stops an order from reaching a supplier.
Data never arrived, yet the interface keeps showing yesterday's figures as today's. That is an outage even though nothing crashed: a decision is being made on a picture that does not exist.
The file went out empty, or to the wrong recipient. The system counts the send as successful, the supplier ships nothing, and the network finds out at an empty shelf.
The web app or the messenger bot will not open. The most visible kind of failure and, oddly, the least dangerous: it produces no wrong decisions.
An error in the data or in a rule that systematically inflates or shrinks orders. Formally the system works — by consequence it is an outage, and it is handled as one.
What is caught without us
These checks came out of specific outages and run on every calculation.
Three hours after sending, the system re-checks: the file was not empty, the recipient exists, a reply is there. The result shows in the order history.
If the data snapshot holds no rows for a warehouse, the send is stopped and the owner gets an explanation instead of a file.
Every calculation carries the date of the data it used. A gap against the schedule triggers an investigation, not a log line.
How fast we respond
The team is small, and promising round-the-clock cover would be a lie. A night-time failure is handled in the morning — before anyone acts on it or an order goes out.
The first move is to switch off whatever can leave the building: order sending on the affected channels. The cause is found afterwards, not instead.
If a network needs response-time commitments, they are fixed by contract with specific figures. Those figures are not on this site, because here they attach to nothing.
What happens afterwards
The affected client gets the cause and the window: which calculations and which sends were touched. No wording along the lines of “degraded service was observed”.
A fix counts not when the failure is gone but when a check exists that will catch its return. The three-hour send verification and the empty-warehouse refusal came about exactly this way.
A change born from an outage lands in the changelog together with its cause. It is the only way to show that the post-mortem was more than talk.
What this page does not have
- There are no uptime percentages here: we do not compute them publicly, and computing them just for a web page would mean inventing them.
- There is no automated status page and no subscription to it. If one appears, it will be in the changelog.
- There is no public incident history: a post-mortem goes to the affected client, not to a public feed.
Show us one critical process. We will show how it runs here.
We look at your cycle: how an order is assembled today, who decides, where time leaks and what the system takes over.