Guide · Updated September 29, 2026
DevOps for startups without a DevOps team
A startup without a DevOps engineer needs two things: a short setup done once, and a weekly check that one person owns. Below are four setup items and five weekly checks, with the GitHub or AWS tool that covers each.
On this page
What DevOps means for a small team
Two jobs matter at this stage. Code reaches production through one repeatable pipeline, and someone notices when CI, cloud cost, runtime health, security alerts or dependencies start to get worse.
Our advice for a company with 2 to 15 engineers is to skip these until a specific problem calls for them:
- Kubernetes, unless you already run it. ECS on Fargate or Lambda needs less upkeep for a few services.
- Multi-region setups. One region with backups and a tested restore covers most early products.
- A service mesh. With a handful of services, load balancers and security groups are enough.
- An internal developer platform. A README and one shared workflow file do the same job at this size.
Set up these four things once
Run CI on every pull request
Add a GitHub Actions workflow that runs on pull_request and runs your tests, linter and build. Then add a branch protection rule or ruleset on main that requires those checks to pass before merging. A failing build then can't reach the deploy step.
Cache dependencies to keep runs short. actions/setup-node, actions/setup-python and actions/setup-go take a cache input, and actions/cache covers the rest. Our suggested target is under 10 minutes for the checks a pull request waits on. Past that, split slow suites into parallel jobs or move them to a nightly run.
Deploy only from the pipeline
Deploy from a workflow that runs on push to main or on a version tag, never from a laptop. Every production change then has a commit, a run log and an author.
Give that workflow AWS access through GitHub's OpenID Connect provider instead of long-lived access keys. The job asks for a short-lived token (the id-token: write permission), and aws-actions/configure-aws-credentials uses it to assume an IAM role scoped to deploys. GitHub's guide Configuring OpenID Connect in Amazon Web Services covers the trust policy. When it works, delete the old access keys.
Add a production environment with required reviewers, so a person approves each deploy. On GitHub Free, Pro and Team plans, required reviewers only work in public repositories (GitHub docs). For a private repository on those plans, require an approved review on pull requests to main instead.
Add alarms before you need them
Create CloudWatch alarms on the metrics that show a service is down or close to it:
StatusCheckFailedon each EC2 instance, alarming on any failed check.ErrorsandThrottleson each Lambda function, alarming above zero or above its normal level.FreeStorageSpaceandCPUUtilizationon each RDS instance: for example, free storage below 10% of what you allocated, or CPU above 80% for 15 minutes.
Add one uptime check that requests your public URL from outside your VPC. A Route 53 health check with a CloudWatch alarm works, and so does any external uptime service. It catches failures your metrics can miss, such as an expired certificate or a DNS mistake.
Send every alarm to one Slack channel through an SNS topic and Amazon Q Developer in chat applications (formerly AWS Chatbot), or to a shared email list. Then create one monthly AWS Budgets cost budget with alerts on actual and forecasted spend. Budgets without actions are free (AWS Budgets pricing). The AWS cost monitoring guide covers the rest of the cost setup.
Keep secrets out of the repository
Turn on GitHub secret scanning and push protection. Both are on by default in public repositories. Private repositories need GitHub Secret Protection, which requires a Team or Enterprise plan (GitHub docs).
Store runtime secrets in AWS Secrets Manager or SSM Parameter Store, and CI secrets in GitHub Actions secrets. Add .env to .gitignore on day one.
If a credential was ever committed, rotate it. Deleting the commit is not enough, because clones, forks and old CI logs can still hold it.
| Item | Tool | Done when |
|---|---|---|
| CI on pull requests | GitHub Actions, a ruleset on main | A pull request with a failing check can't merge |
| Deploys from the pipeline | GitHub OIDC, an IAM role | No AWS access keys are stored in GitHub |
| Alarms | CloudWatch, an uptime check, AWS Budgets | A test alarm (set-alarm-state) reaches Slack |
| Secrets | Secret scanning, Secrets Manager or SSM | Push protection is on and no secret alerts are open |
Check these five things every week
Alarms catch outages. The weekly check catches slow changes that no single alarm shows, such as CI getting slower or one service costing more each week. Compare the last seven days with the seven before. Cost, security and dependencies have longer guides in the guides index.
CI health
Open the Actions tab, filter runs to main, and look at four things:
- Failure rate on
main. A redmainblocks every deploy behind it. - Reruns. A run with
run_attemptabove 1 was retried, usually after a test failed and then passed. It is the simplest flaky-test signal in Actions data. - Queue time: the gap between a run's
created_atandrun_started_at. Long queues point to too few runners or too many parallel jobs. - Duration: the median run time this week against last week.
Suggested warning thresholds, as starting points to tune: success on main below 90% for the week, median duration up 20% on the previous week, or any run queued for more than 5 minutes.
Cloud cost
In Cost Explorer, show the last four weeks at daily granularity, grouped by Service. Compare this week with last week and sort by the change, not the total. For each service that moved, group by usage type to see what changed.
Then read the week's Cost Anomaly Detection alerts. For Cost Explorer users enabled on or after March 27, 2023, AWS creates a monitor for AWS services and a daily email summary by default, at no extra cost (AWS announcement). If your account is older, create that monitor yourself.
Runtime health
In CloudWatch, look at the same metrics your alarms watch, across the whole week:
- EC2: any
StatusCheckFaileddata points, and CPU that stays high. - Lambda:
Errors,ThrottlesandDurationper function. Rising duration adds cost and moves a function closer to its timeout. - RDS:
CPUUtilization,DatabaseConnectionsandFreeStorageSpace. A steady fall in free storage tells you roughly when you will run out.
Treat a missing metric as unknown, not healthy. A Lambda function with zero invocations this week may have a broken trigger rather than a quiet week. An alarm in the INSUFFICIENT_DATA state has nothing to judge, so check whether its resource is idle, deleted or misnamed.
Security alerts
In each repository's Security tab, filter Dependabot and code scanning alerts to those opened in the last week. In AWS Security Hub, filter findings to the CRITICAL and HIGH severity labels. Give each new alert one of three outcomes: fix now, schedule with a date, or dismiss with a written reason.
Dependabot alerts are available on every GitHub plan. Code scanning on private repositories needs GitHub Code Security (GitHub docs). The security alert triage guide covers the order to sort alerts in.
Dependencies and deprecations
Search last week's CI logs for deprecation warnings, such as npm warn deprecated or a Python DeprecationWarning. Then check the end-of-life dates of the runtimes you deploy.
For example, AWS lists April 30, 2027 as the deprecation date for the Lambda nodejs22.x runtime (AWS Lambda runtime policy). AWS emails the account's primary contact at least 180 days before a runtime is deprecated, so make sure that address is an inbox someone reads. The dependency breaking changes guide covers where other notices appear.
| Area | Where to look | Warning sign |
|---|---|---|
| CI | Actions tab, filtered to main | Success below 90%, reruns, queues over 5 minutes |
| Cloud cost | Cost Explorer, grouped by service | A service up on last week with no known cause |
| Runtime | CloudWatch metrics and alarms | Errors, throttles, falling free storage, missing data |
| Security | Security tab, Security Hub | A new CRITICAL or HIGH alert with no owner |
| Dependencies | CI logs, runtime end-of-life dates | A new deprecation, or an end-of-life date within 90 days |
Who owns DevOps when nobody's job is ops
Name one owner each week and rotate the job among engineers. On a fixed day, the owner spends about 30 minutes (our suggestion) on the five checks and writes down each problem with a decision, an owner and a date.
A COO, if you have one, usually doesn't take this over, because production systems stay with engineering. What a startup COO does covers how the COO, CTO and VP of Engineering roles split.
Keep decisions in one place, such as a pinned doc, an issue label or a Slack thread with a fixed name. The next owner starts by reading last week's open items. The weekly operations review guide has an agenda and a decision log to copy.
A rotation spreads the knowledge, so no single engineer is the only one who understands deploys and alarms. It stops working when the checks take several hours a week, or when you need an on-call rotation with one accountable owner. At that point, read when to hire your first DevOps engineer.
Where Grant fits, and what he doesn't do




Grant and his team check the same five areas on a schedule. Maya reads GitHub Actions runs, and Liz reads Cost Explorer cost by service. Noah reads hourly CloudWatch metrics for the EC2 instances, Lambda functions and RDS databases you select. Priya reads existing Dependabot, code scanning and Security Hub findings. David and Owen check your repository against the breaking changes and deprecations Since.dev records from public sources.
On paid plans, check-ins run daily on Pro, every six hours on Studio, and when GitHub reports a workflow run, a check, a push or a new alert. Answers arrive in Slack, the dashboard and MCP, plus an optional weekly digest.
Grant doesn't set up CI, write Terraform or other infrastructure as code, run uptime checks, act as an on-call pager, rerun workflows or change cloud resources. Supported compatibility repairs arrive as pull requests for your review (draft PRs where GitHub offers them).
Grant and his team are AI agents. Their check-ins only read the sources you connect, except Vera's tests of your own agents, and nothing merges or deploys. Meet the team or read about the read-only AWS role.
Questions
How long should a weekly DevOps check take?
About 30 minutes, once the alarms and the budget alert are in place. That is our suggestion, not a measured benchmark. The four setup items are a one-off job, so do them first.
Can one engineer own DevOps part of the time?
Yes, if the checks are written down and someone else can cover when that engineer is away. The risk is one person holding all the knowledge about deploys, alarms and access. A written checklist and a second engineer who runs the check once a month reduce that risk.
What is the least monitoring a new AWS app needs?
CloudWatch alarms on EC2 status checks, Lambda errors and throttles, and RDS free storage and CPU; one uptime check on the public URL; and one AWS Budgets alert. Send all of them to one channel that someone reads.
When should we hire a DevOps engineer?
When infrastructure work keeps pulling engineers off product work, or when uptime, security or a customer contract needs one accountable owner. When to hire your first DevOps engineer covers the signs and the options.
Related guides
- Weekly operations review: a template for small software teamsA 30-minute weekly operations review for small software teams: the agenda, the checks for CI, AWS, security and dependencies, and a decision log to copy.
- AWS cost monitoring for startups: a 15-minute weekly reviewTurn on free AWS cost alerts, read Cost Explorer by service each week and check the common leaks: NAT gateways, idle volumes, log retention and idle RDS.
- When to hire your first DevOps engineerThe signs you need a DevOps engineer, what the first one should own, and how a full-time hire compares with a contractor, a managed platform or software.