Skip to main content
ANCVEIRS
Professional workGuideCloud & Architecture

When a Digital Service Cannot Operate Normally: Designing a Safe Degraded Mode

Resilience is not only recovery time. Services need explicit rules for reduced operation, temporary processing, reconciliation and return to normal.
Zigmārs AncveirsTechnology Leader in FinTech & RegTech · Cybersecurity & Ethical HackingPublished: 26 September 2026Reviewed: 26 September 202614 min

Introduction

A service can have a tested backup, a documented recovery time objective and a competent incident-response team and still have no good answer to a simple failure.

What happens while one critical dependency is unavailable?

Suppose the application is healthy but the identity service is down. Or a national register cannot be queried. Or the payment provider is unavailable. Or a cloud control plane is degraded while the service itself is still reachable.

The usual continuity metrics remain important. RTO tells us how quickly a service or function should be restored. RPO constrains the acceptable data-loss window. Maximum tolerable downtime tells us how long disruption can be endured.

None of them, by themselves, tells the service what it is allowed to do in the meantime.

That gap is the subject of this article.

In August 2026 I submitted a proposal to Latvia’s Crisis Management Centre and Ministry of Smart Administration and Regional Development for a limited public-sector pilot of a methodology covering degraded operation, minimum service and controlled return to normal.1 The proposal did not seek a new legal category or another business-continuity plan. Its question was narrower: can a digital service deliberately move from normal operation into a smaller but trustworthy operating state when a critical dependency is lost, and can it later reconcile temporary activity safely?

Latvia is useful as a case study, but the problem is not Latvian. It sits between service architecture, cybersecurity, business continuity and operational governance across Europe.

Recovery objectives do not define behaviour during the disruption

RTO, RPO and similar continuity metrics answer necessary questions. They do not describe the operating policy of a partially functioning service.

A service with a four-hour RTO may still need to do something useful during those four hours. In another service, continuing any transactional function without a missing dependency may be worse than stopping.

The correct answer is service-specific.

That is why degraded operation should be designed before the incident rather than discovered during it.

Recovery objectives do not define behaviour during the disruption
ConceptPrimary question
RPOHow much data loss can be tolerated?
RTOHow quickly must the function be restored?
Maximum tolerable downtimeHow long can the disruption be endured?
Degraded modeWhich functions may continue, under which controls, until normal operation is safe again?

“Available” and “unavailable” are not the only useful service states

Infrastructure monitoring tends to favour binary language: up or down.

A public-facing or regulated service often needs a richer model. In the 2026 Latvian proposal I used six operating states as a design tool.1 They are not legal categories and should not be presented as an EU standard.

The exact number of modes is less important than three properties: entry criteria, permitted behaviour and exit criteria.

Without those, “degraded mode” is simply incident-time improvisation.

“Available” and “unavailable” are not the only useful service states
ModePurpose
Normalfull intended functionality and primary dependencies
Degradedreduced capability, reduced capacity or an alternative dependency
Minimum serviceonly the predefined essential functions remain available
Offline / manuala bounded non-primary process is used where legally and operationally permissible
Reconciliationtemporary records and actions are compared with the restored authoritative state
Recovery / returnvalidated, controlled transition back to normal operation

Minimum service is not a percentage

A minimum service should not be defined as “keep 30% of the features”.

The starting point is the purpose and legal nature of the service.

Which actions must remain available? Which can wait? Which depend on an authoritative identity, register, payment or trust service? Which functions become unsafe if data freshness is uncertain? Which actions have legal effect that cannot be recreated casually through a manual workaround?

A useful minimum-service definition therefore has at least three classes:

continue — the function can safely operate under predefined controls;

defer — input may be accepted or queued, but the authoritative action is delayed;

stop — the function cannot be performed safely or lawfully without the missing control.

Latvian law already contains a narrower precedent. Cabinet Regulation No. 397 requires ICT critical infrastructure, in the specific context of national threat planning, to identify critical services and the minimum level that must be maintained.2 That does not create a universal obligation for every public information system. It demonstrates that minimum-service thinking already exists within a defined legal scope.

A separate Latvian critical-infrastructure framework uses similar continuity concepts for another legal category while expressly excluding ICT critical infrastructure from its scope.3 The lesson is methodological, not that these regimes can be merged.

Design around service dependencies, not only applications

A digital service is rarely just an application and a database.

Its critical path may depend on identity, authorisation, authoritative registries, APIs, DNS, telecommunications, cloud services, payment rails, timestamping or trust services, messaging providers, supplier support and privileged staff access.

An application can be technically healthy while the service is operationally unusable.

For each material function, ask:

What does this function depend on?

What happens when that dependency is unavailable, delayed or untrustworthy?

Is there an approved alternative?

What must happen to the temporary state when the dependency returns?

The fourth question is where many continuity designs become vague.

Manual fallback is not permission to suspend controls

“Operate manually” can be a legitimate continuity strategy. It can also create a second, poorly controlled information system made of spreadsheets, paper, shared accounts and assumptions.

A manual or offline process should preserve the controls that still matter.

If normal processing verifies identity, representation authority, signatures, account status or an authoritative registry entry, loss of the automated check does not automatically make the check optional.

The service needs an explicit policy for what can be done through an alternative verification method, what can be accepted but held pending verification, and what must stop.

Data protection has the same problem. Temporary local data is still data. Confidentiality, integrity, access control, minimisation and retention do not disappear because the primary platform is unavailable.4

A continuity mechanism that restores availability by silently discarding the controls that make the transaction trustworthy is not a safe degraded mode.

Read-only can be a strong resilience state

Continuity discussions often jump from “online” to “manual processing”.

A read-only state may be safer and more useful.

If a trusted local copy of information remains available, some services can continue to provide information while clearly communicating its freshness boundary. Other services cannot safely do even that because stale information creates unacceptable consequences.

An advisory-only mode can also be useful: the system may continue to guide a user or operator without making an authoritative decision.

These options are deliberately narrower than full service. That is the point.

Degraded operation is not about preserving every feature. It is about preserving the functions whose value still exceeds their risk under the changed conditions.

Reconciliation deserves its own state

Consider a two-hour outage of an authoritative dependency.

During that period a service accepts a bounded set of transactions locally. When connectivity returns, there are now two histories: the authoritative system and the temporary record.

“Synchronise them” is not a sufficient control.

The service may need to detect duplicate execution, concurrent conflicting changes, uncertain ordering, expired authority, missing signatures, stale balances or transactions whose external result is unknown.

Reconciliation should answer:

  • which temporary actions are already reflected in the authoritative state;
  • which remain valid and can be committed;
  • which conflict with later authoritative changes;
  • which outcomes are unknown;
  • which cases require human review;
  • what happens to the temporary evidence after the process completes.

This is why I treat reconciliation as an operating mode rather than a final line in a recovery script.

The risk is especially obvious in payments and ledgers, but it exists in registries, permits, benefits, case-management systems and many other public or regulated workflows.

Recovery and return to normal are not the same event

A dependency responding to health checks does not mean the service is ready for unrestricted operation.

For the categories of NIS2 entities within its defined scope, Commission Implementing Regulation (EU) 2024/2690 requires business-continuity and disaster-recovery plans to address activation and deactivation conditions, recovery sequencing, required resources and restoration from temporary measures.5

The Regulation is not a universal rulebook for every European service. The control logic is nevertheless useful.

Before leaving degraded operation, an organisation should know whether:

  • the original cause has been resolved;
  • the restored dependency is stable;
  • data integrity has been checked;
  • temporary and queued activity has been reconciled;
  • unknown transaction outcomes have been resolved or isolated;
  • emergency credentials and temporary access have been revoked or normalised;
  • monitoring supports a controlled increase in workload.

Return to normal is therefore a governance decision supported by technical evidence, not just a green dashboard.

The 2026 EU resilience guidance moves in the same direction

The European Commission’s July 2026 guidance on the Critical Entities Resilience Directive provides a useful comparison within its own legal scope.6

It encourages critical entities to identify the sub-services that should be restored first, preserve a manual capability when automated or AI-driven systems fail or are tampered with, base continuity planning on business impact analysis, and use phased restart protocols with validation of data integrity and functionality before normal operation is restored.

These are non-binding guidelines for the CER framework, not a general degraded-mode law for every digital service.

Their value here is conceptual. European resilience policy is increasingly interested not only in whether an organisation can recover, but in how it controls temporary operation and resumption.

A dependency-loss scenario

Imagine a Monday morning service outage.

At 09:20 a public digital service loses access to a central electronic-identity dependency. The web application and its database remain available. Existing sessions still work, but new identity assertions cannot be validated.

An improvised organisation has one incident channel and a technical objective: restore connectivity. Individual departments decide for themselves whether to keep serving users.

A service with an explicit degraded-mode design behaves differently.

The authorised owner declares the predefined mode. Existing authenticated users may continue only with functions that do not require fresh identity or representation checks. Transactions requiring new authority are stopped or placed in a controlled pending state where the legal framework permits it. Users see a clear limitation notice. Temporary actions receive unique identifiers and evidence. Engineering continues to restore the dependency, but operational decisions no longer depend on engineering improvisation.

When the identity service returns, full functionality is not immediately restored. Pending actions are checked, conflicts are resolved and the system state is validated first.

There are many possible implementations. The important distinction is between designed degradation and ad hoc exception handling.

A one-page service continuity card

The Latvian proposal suggested a small operational artefact rather than a new parallel continuity-document regime.1

A service continuity card can point to existing BCP, risk and technical documentation while recording the decisions needed during a dependency failure.

The format is not the important part.

Its purpose is to force decisions to be made before the conditions are degraded and information is incomplete.

A one-page service continuity card
FieldDecision to capture
Service and accountable partywho owns the service-level decision
Legal/functional minimumwhat must continue, what may wait, what must stop
Critical dependenciesidentity, data, APIs, network, infrastructure, payments, suppliers
Entry criteriameasurable trigger and authority to declare the mode
Permitted/prohibited actionsthe actual operating boundary
Alternative pathapproved manual, local or secondary process
Data and identity controlshow trust and evidence are preserved
Reconciliationduplicate, conflict and unknown-outcome rules
Exit criteriaevidence required before return to normal
Last exercisewhen this mode was tested and what changed afterwards

Latvia as a practical policy case

Latvia already has several relevant pieces rather than a blank policy landscape.

The National Security Law gives the Crisis Management Centre responsibilities that include coordinating sector continuity and resilience plans, supporting crisis planning, coordinating national crisis exercises and coordinating follow-up evaluation after crises and exercises.7

Cabinet Regulation No. 397 sets minimum cybersecurity requirements for the entities within its scope and, for ICT critical infrastructure in national-threat planning, requires identification of critical services and their minimum level.2

Cabinet Regulation No. 399 defines the holder of a state administration service as the institution or other legal subject competent to ensure that service.8 That matters because a degraded-mode methodology cannot move responsibility to a central technology authority: the service owner or other legally competent body still has to determine the lawful and operational minimum.

My August 2026 proposal therefore suggested a pilot, not a new universal mandate: select a small number of services with different dependency patterns, create a concise continuity card, exercise loss of a shared or external dependency, require a reconciliation phase, and decide only afterwards whether common guidance adds value.1

That remains the useful sequence. Map, exercise, measure, then regulate only if a real gap remains.

The European context

NIS2 includes business continuity, backup management, disaster recovery and crisis management in its cybersecurity risk-management measures.9

Commission Implementing Regulation (EU) 2024/2690 provides more detailed operational requirements for specific categories of NIS2 entities, including recovery order and resumption from temporary measures.5

The CER framework addresses a different legal population and a broader all-hazards resilience problem, but its 2026 guidance also emphasises service prioritisation, manual capability and validated phased restart.6

None of these sources defines the six-state model used in this article.

What they collectively support is a narrower proposition: continuity planning must deal with more than the final recovery deadline. It must deal with the behaviour of the service during disruption and the controlled transition out of temporary measures.

How to test whether the design is real

A tabletop exercise is more revealing than another policy review.

Choose one service with meaningful dependencies. Remove one of them in the scenario. Do not let the team assume that it will return in ten minutes.

Then ask who can change the service mode, what functions remain available, how users are told, how temporary actions are identified, what happens after four hours, what happens if the dependency returns only intermittently, how conflicting records are handled and who authorises the final return to normal.

If most answers begin with “we would decide during the incident”, the organisation may have recovery documentation but it does not yet have an executable degraded mode.

Conclusion

Continuity architecture is often drawn as a jump from normal operation to recovery.

The difficult part is the space between them.

That is where a service must decide which functions remain trustworthy, which controls cannot be waived, what can be deferred, how temporary activity is recorded and how it will later be reconciled with the authoritative state.

A safe degraded mode is therefore not “keep as much running as possible”.

It is a predefined operating boundary: what may continue, what must stop, who can decide, what evidence must be preserved, and what must be true before normal operation resumes.

Frequently asked questions

Is “degraded mode” an EU legal term for digital services?

Not in the general sense used in this article. Here it is an engineering and governance term for controlled reduced operation. Specific laws and sectoral rules may define their own continuity states and obligations.

Does every digital service need a minimum-service mode?

No universal rule of that kind is claimed here. For some services the safest degraded state is complete stop. The appropriate design depends on service purpose, risk, dependencies and applicable law.

Can an organisation bypass identity controls during an outage?

An outage does not itself remove identity, authorisation, data-protection or legal-validity requirements. Any alternative process needs its own lawful and secure basis.

Why is reconciliation separate from recovery?

Because restoring a dependency does not resolve temporary records, duplicate actions, conflicting updates or unknown transaction outcomes. Those need an explicit process before the service state can again be treated as authoritative.

Is a manual process always a good resilience measure?

No. Manual fallback is useful only when its authority, scope, evidence, data handling and reconciliation are defined. An uncontrolled spreadsheet or shared credential can create a larger incident than the original outage.

Is this a replacement for BCP or disaster recovery?

No. It is an operational layer within or alongside existing continuity planning. BCP and disaster recovery remain the broader governance structures.

Source status

Legal and policy sources were checked on 25 September 2026. The six operating modes and service-continuity-card structure are the author’s proposed methodology, not statutory requirements.

This article is an analysis of technical, organisational and public-governance practice, not individual legal advice.

Sources

  1. Zigmārs Ancveirs, “Par valsts digitālo pakalpojumu ierobežotas darbības režīmu un minimālā pakalpojuma apmēra metodikas izvērtēšanu un pilotēšanu”, 24 August 2026. Author’s document archive.
  2. Latvia, Cabinet Regulation No. 397 of 25 June 2025, “Minimālās kiberdrošības prasības”, including Annex 9 · Likumi.lv
  3. Latvia, Cabinet Regulation No. 10 of 13 January 2026, “Kritiskās infrastruktūras apzināšanas, darbības nepārtrauktības, drošības un noturības pasākumu plānošanas, īstenošanas un incidentu paziņošanas kārtība”,… · Likumi.lv

    Latvia, Cabinet Regulation No. 10 of 13 January 2026, “Kritiskās infrastruktūras apzināšanas, darbības nepārtrauktības, drošības un noturības pasākumu plānošanas, īstenošanas un incidentu paziņošanas kārtība”, particularly Section 4 and Annex 1

  4. Regulation (EU) 2016/679 (GDPR), particularly Articles 5 and 32 · EUR-Lex
  5. Commission Implementing Regulation (EU) 2024/2690, Section 4, within its defined entity scope · EUR-Lex
  6. European Commission, “Guidelines on the application of Article 13(5) of Directive (EU) 2022/2557 on the resilience of critical entities”, C/2026/3712, 13 July 2026, particularly paragraphs 66–72 · EUR-Lex
  7. Latvia, National Security Law, Section 23.^1 · Likumi.lv
  8. Latvia, Cabinet Regulation No. 399 of 4 July 2017, “Valsts pārvaldes pakalpojumu uzskaites, kvalitātes kontroles un sniegšanas kārtība” · Likumi.lv
  9. Directive (EU) 2022/2555 (NIS2), Article 21 · EUR-Lex