CertClue
Courses · Cloud
Junior SRE / Linux Admin

Linux & Site Reliability Foundations

Learn to keep a Linux service up: read a machine under load, choose what to measure, define an objective you can defend, take a page at 3am without making it worse, and write the postmortem that stops it happening again.
6 hrs taught · 6 to 11 hrs applied 7 modules 29 lessons 7 portfolio artifacts Completion certificate Updated August 2026
Created by the CertClue team
What you'll build

Real portfolio pieces built during the course, not a certificate for its own sake. Each one is work you can show.

  • Monitoring dashboard specA one screen dashboard specified panel by panel: the service level indicator and budget on the top row, the four golden signals below it, the dependencies that explain them, and a written note of what you deliberately left off and why.
  • On call runbookA runbook for the database timeout alert, written for a tired stranger: what it means in user terms, five read only checks with expected values, what to capture before mitigating, when to escalate, and the do not list.
  • Automation scriptA disk pressure triage script that is read only by default, checks its assumptions and refuses clearly when they do not hold, is safe to run twice, and hands the destructive options back to a human.
  • Blameless postmortemA write up of the 3am incident with impact in user terms, a timeline built from captured evidence, contributing factors that name mechanisms rather than people, what went well, where you got lucky, and actions with owners and dates.
  • SLO and error budgetOne page defining the indicator as a ratio of good events to valid events, the exclusions and why each is fair, the target with the four inputs behind it, the budget arithmetic, two burn rate alerts and the agreed policy for an exhausted budget.
  • Alert reviewEvery alert that woke somebody this week, reviewed the way you would diagnose a service: what it actually measures against what it is trying to protect, its own firing history with the actions recorded beside it, a verdict of rewrite, delete or keep with what settled it, the replacement written out as an expression rather than as a philosophy, and the detection gap the change creates stated out loud rather than glossed over.
  • Change plan with a rollbackOne risky change planned so that its cost lands where it does not matter: the window and why that window, the steps in the order they happen with what has to be true before each one starts, what it costs while it runs and out of whose budget, the number that stops the work with the person who can call it, and the rollback written before anything is touched. Including an honest answer to whether it should happen at all this week.

What you'll learn

Fundamentals, kept short
The role decoded
A day in the role
The recurring calendar
Live simulation · a worked on call week
Artifacts · the portfolio you leave with
Handoff

Course content · 7 modules, 29 lessons

Sign up to unlock every lesson - the titles below show exactly what is inside.

Enough Linux to reason about a running system, enough networking to answer the most common page there is, and enough reliability theory to know what to measure and what to promise. Eight lessons, then you are working.

What reliability work actually is
Linux as a system you can reason about
systemd units and the journal
Reading a machine under load
Memory, disk, and the numbers that mislead
When you cannot reach it
The four golden signals, and alerting on symptoms
SLIs, SLOs and the error budget

Requirements

  • Comfortable with the fundamentals this course's own Module 1 covers, or equivalent experience.
  • No prior experience in this field is required to start.
  • A computer with a reliable internet connection.
  • Comfortable using a web browser - no software to install.

Description

Every CertClue course follows the same seven-part shape: fundamentals, the role translated out of job-posting language, a real working day, the job's recurring rhythms, a multi-day simulation, the portfolio you build along the way, and a handoff into your next move. Here is what that looks like for junior sre / linux admin.

Who this course is for

Anyone aiming to become a junior sre / linux admin, including career changers with no background in it yet. This is the entry rung of a realistic ladder:

entry
Junior Linux Administrator

Maintains Linux servers, handles routine maintenance, and responds to basic alerts.

Server maintenanceAlert responseScripting basics
mid
Site Reliability Engineer

Owns uptime for a set of services, builds monitoring and automation, and leads incident response.

MonitoringAutomationIncident response
senior
Senior SRE

Sets reliability standards across the org, designs for scale, and mentors on-call engineers.

Reliability engineeringScale designOn-call mentoring

Where it leads

This course prepares you for the CompTIA Linux+ then RHCSA role or credential path. Named for preparation only - no partnership or endorsement is implied.

Reviews

No reviews yet. Reviews come from learners who have taken the course, so this stays empty until someone leaves one.

Sign in to leave a review.

Students also explore