Skip to content

Legal

Service Level Agreement

Version
0.1
Status
In legal review
Last updated
21 September 2026

In legal review.

The technical commitments below are derived from how the service is built. Our lawyers are reviewing the contractual framing, including the remedies and how this agreement works with the Terms of Service, and it may change before it takes effect.

1. What this document covers

This Service Level Agreement applies to production Balta DB database services. It sets out the availability we commit to, what we do when we miss it, and — equally important — what we do not promise.

Two commitments, and no others. This agreement commits to monthly availability (section 3) and to a recovery point objective (section 4). Everything else here describes how we work: what we do when we miss a commitment, how maintenance is announced, how incidents are communicated, and how support operates. Nothing else in this document is a target you can hold us to, and we say so rather than leaving you to infer it from which sentences happen to carry a number.

We have written it to be read before you buy, not after something goes wrong.

2. Architecture, stated first

Balta DB services run on a single node. One PostgreSQL instance on one host, with no standby and no automatic failover.

Every commitment in this document is a single-node commitment. Nothing in Balta DB, our marketing,

or this agreement should be read as promising high availability, replication or automatic failover.

3. Availability commitment

We commit to 99.5% monthly availability, measured per service, excluding announced planned maintenance.

That permits approximately 3 hours and 39 minutes of unplanned unavailability per calendar month.

We publish the arithmetic behind that figure because it is the honest one for this architecture: a single host loss requiring a rebuild from backup consumes a meaningful part of that budget for a mid-sized database. A 99.9% commitment — under 44 minutes — could not survive one such event, and committing to it would be committing to something the architecture cannot deliver.

3.1 How availability is measured

A service is available when it accepts a TCP connection, completes a TLS handshake, authenticates, and answers a trivial query.

Measurement is taken from at least two external probe locations on different networks, at 30-second intervals. A service is counted unavailable only when a majority of probes fail, and a window in which fewer than two independent networks reported is recorded as uncovered rather than as available.

For any period that no external measurement record covers, a credit claim under section 6 is assessed from our own incident record and from the evidence you supply. The probe locations, and the date external measurement began, are recorded in the changelog.

3.2 What is excluded

  • Announced planned maintenance within its window and duration (section 5)
  • Downtime caused by your own configuration, queries, workload or connection exhaustion
  • Suspension for non-payment or breach of the Acceptable Use Policy
  • Force majeure, as defined in the Terms of Service.

Location-level provider outages are not excluded. If the datacentre or hosting provider we chose fails, that is our supplier and our decision, not your problem to absorb. Credits apply.

4. Recovery objectives

4.1 Recovery point objective — how much data you could lose

We commit to an RPO of 5 minutes or better under normal operation.

RPO here means the gap between now and recoverable_to — the latest point in time we can actually restore you to. That requires a base backup which has passed verification and an unbroken chain of write-ahead log segments following it.

It is not a configuration value we quote at you, and it is not simply "how recently we uploaded something". A recent segment proves nothing if an earlier one is missing: a gap makes every point after it unreachable. We check the chain continuously, we alert on a gap immediately, and we show the result per service in your dashboard — separately for each of the two repositories. If we have not been able to evaluate the chain recently, your dashboard says so rather than showing you a stale green tick.

FailureData loss
PostgreSQL crash, host reboot, or host OS crash with storage intactNone. Unarchived write-ahead log is still on local storage and is replayed
Host lost including its storageLimited to unarchived write-ahead log: the RPO exposure shown at the time of failure

Because local storage is RAID-protected, the total-storage-loss case is the less likely of the two.

This commitment is suspended for any period during which we have alerted you to a failure in the archive pipeline for your service. We are obliged to raise that alert promptly. We can commit to an RPO only while the pipeline we monitor is working, and we commit to telling you quickly when it is not.

4.2 Recovery time objective — how long restoration takes

A PostgreSQL process failure, or a host reboot with storage intact, is recovered automatically: PostgreSQL restarts and replays the write-ahead log still on local storage.

A host lost outright is recovered by rebuilding the service from backup on another host in the same location, using capacity held in reserve for exactly that.

The loss of an entire location, provider or datacentre has no committed recovery time. See 4.3.

Recovery times are published by database size band once recovery drills at 10 GB, 100 GB, 500 GB and 1 TB have produced them. Published figures are the measured 95th percentile with margin, and we will not publish an estimate in its place. The time an automatic restart takes is published once it has been measured on production hardware.

Recovery time is measured end to end: detection, decision, target host selection, provisioning, restore, write-ahead log replay, startup, verification and endpoint change. Not restore duration alone.

Routine restore verification does not measure this. A verification run restores one database, uncontended, onto a host that already exists, and it measures none of detection, host selection, provisioning, fencing or the endpoint change. Its duration is a lower bound on recovery, not a recovery time. The published figure comes from scheduled recovery drills, which exercise the whole sequence, and it is refreshed when a drill runs. Treating a verification duration as an RTO would publish a number smaller than the one you would actually experience.

4.3 Loss of an entire location

If a location, provider or datacentre becomes unavailable, services in that location may remain unavailable until the provider recovers, or until our team performs a manual recovery where that is practical.

Balta DB does not provide automatic cross-region failover, active-active regions or cross-region synchronous replication, and does not commit to a recovery time for this scenario.

Your data survives. Backups are held in two repositories, and the second is in a different country. A location loss is an availability event, not a durability event.

Service credits apply to this scenario in the normal way.

5. Planned maintenance

  • Each service has an assigned weekly maintenance window of at least four hours, stated in UTC on the service's Configuration tab. Selecting your own window is not available.
  • Maintenance with expected customer impact is announced at least 72 hours in advance, stating the affected service, the reason, the scheduled time, the expected impact and the expected duration.
  • Maintenance is excluded from the availability calculation only if it occurs within the announced window and does not exceed the announced duration. Overrun counts as unplanned downtime. This is what makes the exclusion honest rather than a loophole.

5.1 Emergency maintenance

Permitted only for an actively exploited security vulnerability, hardware pre-failure indicators, a data-integrity risk, or a provider-driven emergency.

Notice is given as early as the situation allows. Every emergency maintenance event produces a written customer-facing explanation within five business days.

5.2 Network addresses and the notice before they change

Each location publishes the IP set its services live inside, on /locations. A service does not move outside the set published for its location, so a firewall rule written against that list keeps working.

When we add hardware on a new range we publish the range at least 30 days before any existing service migrates onto it. Each published range carries the date it was published, and that is the date the 30 days run from. A rule written against yesterday's list keeps working throughout.

Two limits on that commitment, stated because they are the cases where it would otherwise be read as more than it is:

  • A new service may be placed on a newly published range immediately. You are told your endpoint when it is created, so there is nothing you wrote a rule against yet.
  • In an emergency — the conditions in §5.1 — we may move a service onto a range inside its notice period. That requires a recorded reason, names the services it covers, notifies you, and raises an internal alert. It is not a quiet path, and it is not a general exception.

The 30 days is the same notice we already owe you for a new subprocessor, because a new infrastructure provider is one.

6. Service credits

Monthly availabilityCredit
Below 99.5% and at or above 99.0%10% of that service's monthly fee
Below 99.0% and at or above 95.0%25%
Below 95.0%50%

Claims must be submitted within 30 days of the end of the affected month. Credits are applied to a future invoice, are capped at 100% of the affected service's monthly fee, and are the sole and exclusive remedy for failure to meet the availability commitment.

Claims are assessed and applied by hand. There is no automatic credit calculation, and until the probes in section 3.1 are running there is no external measurement to compute the monthly figure from. Email us and we will show you our working.

7. Incident communication

  • Directly affected customers are notified by email.
  • Every Severity 1 incident receives a written post-incident report within five business days, covering impact, timeline, cause and corrective actions.

8. Support

Infrastructure monitoring and operator on-call run 24/7, independent of any support plan.

Contact support any time by email or from your dashboard.

Every service is monitored around the clock.

This agreement commits to no first-response time. We answer as fast as we can and we will not write down a number we have not measured ourselves against. Severity sets the order we work in, and these are the definitions we use:

SeverityWhat it means
1Service unavailable, data at risk, or security incident
2Service degraded or materially impaired
3Functional problem with a workaround
4Question, request or documentation issue

9. Changes to this agreement

We may update this agreement as the service changes. Changes that materially reduce the commitments here will be notified to account holders by email at least 30 days before taking effect, and the previous version remains available. Every version is listed at the foot of this page.

Version history

Every version of this document stays available. Where a change materially reduces the commitments in it, account holders are notified by email at least 30 days before it takes effect.

  • Version 0.1 · 21 September 2026 · First version, written from our own architecture. In legal review.