Stays Up
Stays Up · infrastructure

Your infrastructure stays up

We help engineering teams build, fix and run reliable cloud infrastructure: architecture, migrations, security, cost and AI reliability.

Written reply within one business day. Fixed prices, invoiced on delivery.

↓ scroll

What we do

Everything a DevOps hire would do, without the hire

Pick one piece or hand over the whole thing. Most teams start with an audit and keep us on for the parts they'd rather never think about again.

Run

Migrations

New environments built properly from the start, or existing ones migrated: between clouds, off bare metal, or away from a setup that grew by accident and nobody fully understands.

CI/CD

Build and deployment pipelines in GitHub Actions, GitLab CI or ArgoCD. Automated tests, staged rollouts, one-click rollback, and deploys that any engineer on the team can run safely.

Kubernetes setup and management

Cluster design, upgrades, autoscaling, GPU node pools, resource limits and cost control on EKS, GKE or self-managed Kubernetes. Also honest advice about whether you need Kubernetes at all.

Infrastructure as code

Your environment described in Terraform or Pulumi: version controlled, reviewable, reproducible. No more configuration that exists only in the console and in one person's memory.

Database reliability and performance

PostgreSQL and MySQL tuning, connection pooling, replication, failover and slow-query work, for teams whose database has quietly become the thing everything waits on.

Backup and disaster recovery

Backups that are verified by actually restoring them, documented recovery procedures, and a tested answer to how long you'd be down in the worst case.

Protect

Security audit and hardening

Cloud security review covering IAM permissions, public exposure, secrets management, unpatched images and MFA gaps, each finding ranked by real risk rather than scanner severity.

SOC 2 and ISO 27001 groundwork

The technical controls auditors ask for, access control, logging, encryption and change management, put in place before the audit rather than during it.

EU AI Act logging readiness

High-risk record-keeping (Article 12) now applies from December 2027, and transparency duties are already live. Six-month retention means the logs must exist well before the deadline. We build it calmly now, not in a 2027 panic.

Measure

Cloud cost optimisation

AWS and GCP bills cut by rightsizing instances, removing idle and orphaned resources, fixing egress and NAT charges, and buying the right commitments. Typically 20–40% off the monthly bill.

Monitoring and observability

Metrics, logs and traces in one place using Prometheus, Grafana and OpenTelemetry, LLM and agent traces included. Alerts tied to things that actually matter, so you hear about problems before your customers do.

AI agent reliability

Your agents report success on every run. We instrument what they actually did, claimed against actual, so a job that loops forever while logging "completed successfully" shows up on a dashboard instead of in an incident.

How it works

Four steps, one week

1 · Call

Tell us what you run

A free 30-minute call, or three lines by email: what you are running, what is bothering you, which audit fits. We confirm the scope and the fixed price in writing.

2 · Audit

One week, read-only

A scoped read-only role on one account, or an export of your traces. We change nothing. What we receive and when it is deleted is on the trace handling page.

3 · Result

The report, and a walkthrough

Findings ranked by what they cost you, each with a reproduction, plus the checks or queries we used. A 30-minute walkthrough of what we would fix first.

4 · Payment

Invoiced on delivery

The fixed price agreed in step one, invoiced when the report lands. Anything you want done after that is quoted as a fixed price before it starts.

Pricing

Clear prices, no lock-in

Start with an infrastructure audit. We spend a week on your cloud, security, reliability, delivery pipeline and costs, then hand you a ranked list of what we'd fix and why. Implement it yourself, give it to someone else, or bring us in to do the work. The report is yours either way.

Start here

Infrastructure audit

€500 fixed
Delivered within a week · read-only access
  • Cost, security, reliability and delivery reviewed
  • Written report, typically 8–15 findings
  • Each finding priced and ranked by what it costs you
  • 30-minute walkthrough call
  • Yours to keep, no obligation
  • One cloud account, one environment
Book by email
 

Custom

Let's talk
Fixed price, quoted before it starts
  • SOC 2 and ISO 27001 groundwork
  • Architecture and redesign
  • Cloud migrations, between providers or off bare metal
  • Infrastructure-as-code rebuilds
  • Multi-environment and multi-region setups
  • Cost optimisation, priced against the savings instead
Tell us what you need

Both audits are a fixed price, invoiced when the report is delivered. Anything beyond an audit is scoped and quoted as a fixed price before it starts. Nothing is billed that was not agreed in writing.

We invoice from Portugal. EU business-to-business services are invoiced under the applicable reverse-charge rules. Bank transfer or card; USDC on request.

Contact

Tell us what you run

Three lines is enough: what you are running, what is bothering you, and which audit you want. We reply in writing within one business day with a price or a straight no.

Send a message

This is the front door. Most work starts and finishes by email.

Or just email hello@staysup.io

Book a call

A free 30-minute review of your setup. Video call, no preparation needed.

Open the calendar

Typically available within 2–3 working days

FAQ

Questions people ask

If yours isn't here, email it. You get a straight answer either way.

Do we really not need a call?

No. Email what you run and which audit you want; we confirm scope and price in writing, we start, and you are invoiced when the report is delivered. Read-only access is arranged by email too. The free review call exists for people who prefer to talk, and skipping it changes nothing about the work.

What does the AI agent reliability audit actually check?

Three things, on your real traces. Whether the agent's completion claims match the state it left behind. Whether it loops, stalls or keeps working after it has lost the thread, and how many steps before anything visible happens. And how its tool use fails: tools skipped, results ignored, outputs fabricated, forbidden state changes reported as success. You get the findings with reproductions, and the detectors stay with you.

Do we have to give you access to our cloud account?

For the audit, read-only is enough: a scoped role with no write permissions, which you can revoke the day the report lands. For work we actually carry out we need write access to the parts we're responsible for, scoped to those. You keep the root account throughout.

We already have someone who handles infrastructure.

Then you probably want the audit, or a fixed-price piece now and then. A second pair of eyes on someone else's work is worth €500 once a year, and we'll tell you honestly if there's nothing to do. Plenty of our work is for teams who already have infrastructure people and need a specialist for one problem.

What if the audit finds nothing serious?

Then you've paid €500 for a clean bill of health and can stop thinking about it. It happens, though not often. Usually the cost findings alone cover the fee within a month or two.

Are we committing to anything ongoing?

No. The audit is a one-off and any further work is quoted as a fixed price per piece. There's no contract to cancel and nothing renews on its own.

What if something breaks at 3am?

We're not a 24/7 NOC and won't pretend otherwise. One small team can't honestly promise that. What we do instead is set things up so 3am incidents mostly stop happening, and so the ones that do have a runbook your team can follow without us.

Which clouds and tools do you work with?

AWS, GCP and Azure; Kubernetes, Terraform, Pulumi, Docker, Helm, Ansible, GitHub Actions, GitLab CI, ArgoCD, Datadog, Grafana, Prometheus, OpenTelemetry, Vault and PostgreSQL. If you're on something unusual, say so on the call and we'll tell you straight whether we're the right people.

Does the EU AI Act apply to us?

If you ship AI features to EU users: transparency obligations (Article 50: telling users they're talking to AI, machine-readable marking of generated content) apply since 2 August 2026. High-risk record-keeping (Articles 12, 19 and 26) was deferred to 2 December 2027 by the Digital Omnibus. But the six-month retention requirement means the logging has to be running well before that date. Whether a given system is high-risk is a question for your lawyer; we build the technical side from their answer, and the same logging is worth having either way.