Skip to content
GrantbySince.dev
GrantWatchesChangesDocsPricing
Sign inStart free
  1. Home
  2. /Guides
  3. /Weekly operations review

Guide · Updated September 29, 2026

Weekly operations review: a template for small software teams

A weekly operations review is a 30-minute meeting where one owner walks through CI, cloud cost, runtime, security and dependencies, and the team records a decision for every open problem. Below are the agenda, the checks and a decision log you can copy.

On this page

  1. What a weekly review catches that alerts miss
  2. Who attends and how to prepare
  3. The agenda
  4. What to check in each area
  5. How to record decisions
  6. A template you can copy
  7. Where Grant fits, and what he doesn't do
  8. Questions
  9. Related guides
On this page
  1. What a weekly review catches that alerts miss
  2. Who attends and how to prepare
  3. The agenda
  4. What to check in each area
  5. How to record decisions
  6. A template you can copy
  7. Where Grant fits, and what he doesn't do
  8. Questions
  9. Related guides

What a weekly review catches that alerts miss

Alerts fire on single events: a failed deploy, a budget threshold, a new critical CVE. Slow changes don't trip them. A weekly look catches these:

  • CI runs that take longer or wait longer in the queue each week. No single run fails, so nothing alerts.
  • Cost in one service growing week on week while staying under your budget alert.
  • An alert backlog that grows because nobody owns the new items.
  • A deprecation date getting closer. AWS emails the account's primary contact at least 180 days before a Lambda runtime is deprecated (AWS runtime policy), and one email to one inbox is easy to miss.

A week also fits the data. Cost Explorer shows cost by day and refreshes at least once every 24 hours (AWS), so one week gives you seven days to compare with the seven before. A week of CI runs on main shows whether a failure was one bad day or keeps coming back.

Who attends and how to prepare

  • The owner prepares and runs the meeting: the CTO, or an engineer on a rotation.
  • One other engineer attends, so at least two people know the current state.
  • A founder joins when a cost or risk decision is expected. Accepting a risk or approving spend is their call.

Use the same day and time every week. An early-week slot leaves the rest of the week to act on the decisions (our suggestion).

Prep takes about 15 minutes once alerts are set up (our estimate). If you run the full AWS cost or security routine from its own guide, it replaces that area's prep. The owner:

  1. Lists recent failed runs on main with gh run list -b main -s failure, or from the Actions tab filtered by branch.
  2. Opens Cost Explorer at daily granularity, grouped by Service, for the last 14 days.
  3. Checks CloudWatch alarms, each repository's Security tab and Security Hub.
  4. Copies last week's open decisions into this week's notes.

If the prep shows nothing new in any area, skip the meeting, send the notes and write “no change” in the log.

Without an operations hire, this meeting is where every open problem gets an owner. What a startup COO does covers who owns operations until then, and when the role needs its own person.

The agenda

Seven items in 30 minutes. The time boxes are a starting point. If an item needs a longer discussion, give it an owner and its own meeting.

1. Open decisions from last week (5 min)

  • Is each decision done, on track or late?
  • For each one marked done, what new evidence shows it: a passing run, a closed alert, a lower daily cost?
  • Has any accepted risk reached its review date?

2. CI and deploys (5 min)

  • Is main green? Which workflows failed this week, and why?
  • Which runs needed a rerun (run_attempt above 1)? A rerun that passes often points to a flaky test.
  • Did run duration or queue time grow against last week? Did any deploy fail or get rolled back?

3. Cloud cost (5 min)

  • Which services changed most against last week? Sort by the change, not the total.
  • Did Cost Anomaly Detection flag anything?
  • Does something you shipped or planned explain each increase?

4. Runtime (5 min)

  • Which CloudWatch alarms fired, and did each one need action?
  • Is any alarm in INSUFFICIENT_DATA, or has a metric stopped reporting? Treat missing data as unknown, not healthy.

5. Security (5 min)

  • Which Dependabot, code scanning and Security Hub items are new at critical or high?
  • Is any new CVE in CISA's Known Exploited Vulnerabilities catalog? A listed CVE has confirmed exploitation, so it goes first whatever its score.
  • Does every open critical item have an owner?

6. Dependencies and deprecations (3 min)

  • Did CI logs show new deprecation warnings, such as npm warn deprecated (older npm versions print it in capitals)?
  • Does a runtime, base image or major dependency reach end of life within 90 days (our suggested window)? Example: AWS lists April 30, 2027 as the deprecation date for the Lambda nodejs22.x runtime (AWS runtime policy).

7. Decisions and owners (2 min)

  • Read back each new decision with its owner and date.
  • Confirm that no open item is left without an owner.

What to check in each area

Each check has a place in GitHub or AWS. Compare with last week, not only the current value.

Weekly checks by area, and where to find each one
AreaCheckWhere
CIFailed runs on maingh run list -b main -s failure
CIRerunsrun_attempt above 1 in the workflow runs API
CIRun durationRun history in the Actions tab
CIQueue timerun_started_at minus created_at
CostChange by serviceCost Explorer, daily, grouped by Service
CostAnomaliesCost Anomaly Detection daily summary
RuntimeEC2 status checksCloudWatch StatusCheckFailed
RuntimeLambda errors and throttlesCloudWatch Errors, Throttles
RuntimeRDS storage and connectionsCloudWatch FreeStorageSpace, DatabaseConnections
SecurityNew dependency alertsDependabot, in the Security tab
SecurityNew code alertsCode scanning, in the Security tab
SecurityAWS findingsSecurity Hub, severity CRITICAL or HIGH
DependenciesDeprecation warningsCI logs and run annotations
DependenciesEnd-of-life datesAWS Health Dashboard, vendor release pages

What to bring to the meeting

These are our starting thresholds. Adjust them after a month of reviews.

  • Any failed run on main without a known cause, and any rerun that passed with no code change.
  • Run duration or queue time up by more than 20% on last week.
  • A service's weekly cost up by more than 20%, or a new service on the bill.
  • Any alarm that fired more than once, and any metric with no data.
  • Any new critical or high alert, and any CVE listed in KEV.

Write down missing data as its own line. A Lambda function with zero invocations may have a broken trigger. A repository with no code scanning alerts may have no scanning set up.

Lambda functions deployed as container images get no runtime deprecation notices from AWS (AWS runtime policy), so check their base images yourself.

Guides for each area: DevOps for startups (CI and runtime), AWS cost monitoring, security alert triage and dependency breaking changes. See all guides.

How to record decisions

Keep one decision log for the team: a shared doc, a spreadsheet or a labelled GitHub issue. Each entry has five fields:

  • Finding: what you saw, in one line.
  • Evidence: a link to the run, the Cost Explorer view or the alert.
  • Decision: one of the four below.
  • Owner: one person, not a team.
  • Date: when it will be done or reviewed again.

Allow four decisions only:

  • Fix now: the owner starts this week.
  • Schedule: fix by a named date.
  • Accept the risk: write the reason and a review date. The item returns to the agenda on that date.
  • Not an issue: write the reason, such as a false positive or a vulnerable package used only in tests.

Two rules keep the log useful. Every open finding has an owner and a date. An item counts as addressed only when there is new evidence (a green run, a closed alert, a lower daily cost), not when someone says it's done.

A template you can copy

Paste this into your notes each week and keep every copy, so you can compare weeks.

Weekly operations review, <date>
Owner: <name>
Attendees: <names>

Open decisions from last week
- <finding>: <done | on track | late>

CI: <green | failures, reruns, slower>
Cost: <top change by service vs last week>
Runtime: <alarms, missing metrics>
Security: <new critical or high, any in KEV>
Dependencies: <deprecations, EOL within 90 days>

Decisions
- <finding> | <decision> | <owner> | <date>

Next review: <date>

Decision values: fix now, schedule, accept risk (with a review date), not an issue (with a reason).

Where Grant fits, and what he doesn't do

Grant, COO

Grant is the COO for your software operations, and his team can do much of the prep for this meeting. On a schedule, Maya reads GitHub Actions runs and Liz reads Cost Explorer by service. Noah reads CloudWatch metrics for the resources you select, and Priya reads existing Dependabot, code scanning and Security Hub findings. David and Owen check compatibility, configuration and lifecycle.

Check-ins run daily on Pro and every six hours on Studio, plus on GitHub events. An optional weekly digest to Slack or a signed webhook links to the latest check-in (Slack and notifications).

In a conversation you can assign an owner, defer a finding to a date or mark it addressed. The next check-in compares new evidence with that decision (consults and decisions). Marking a finding addressed records your decision; it doesn't prove the issue is fixed.

Grant doesn't run the meeting or decide for you. Apart from supported compatibility repair PRs, he doesn't fix what his team finds. Grant and his team are AI agents. Their check-ins only read, except Vera's tests of your own agents, and nothing merges or deploys.

See how follow-through works.

Questions

Should the review be weekly or daily?

Weekly. Urgent problems, such as a failed deploy or a new critical alert, should reach you through alerts on the day they happen. The weekly review is for trends and ownership. Meet daily only for a set period when something calls for it, such as a launch or an open incident.

What if there's nothing to report?

Write “no change” for each area and end early. The log then shows the week was checked. If the prep already shows nothing new, skip the meeting and send the notes.

Who should run it without an ops hire?

The CTO at first, then engineers on a rotation, changing each week or month. A rotation spreads knowledge of CI, AWS and the alert queues, so it doesn't sit with one person. A founder joins when a cost or risk decision is on the table.

How is this different from a sprint retro?

A retro looks at how the team worked. The operations review looks at the state of your systems: CI, cloud cost, runtime, security and dependencies. Keep them as separate meetings so each gets its full time.

Related guides

  • DevOps for startups without a DevOps teamSet up four things once, then check five every week. A practical DevOps checklist for startups on GitHub and AWS that have no DevOps engineer yet.
  • What does a COO do at a startup?What a startup COO owns, when companies add one, how the role differs from CTO and VP of Engineering, and who runs software operations before then.
  • How to triage Dependabot, code scanning and Security Hub alertsA 20-minute weekly routine to triage Dependabot, code scanning and AWS Security Hub alerts: known exploitation first, then exposure, then severity.
Hire Grant free See plans

Grant. Your COO and his ops team.

Product

  • Meet Grant
  • Conversation
  • Compatibility repairs
  • Watches
  • Pricing
  • FAQ

Guides

  • All guides
  • DevOps for startups
  • Startup COO
  • Weekly operations review
  • First DevOps engineer
  • AWS cost monitoring
  • Security alert triage
  • Dependency breaking changes

Developers

  • Docs
  • Quickstart
  • HTTP API
  • MCP server
  • Webhooks
  • llms.txt

Public data

  • Changes
  • Sources
  • Atom feed

Company

  • Privacy
  • Terms
  • Contact
  • Security
since.dev

© 2026 Since.dev