Under VIQ7, an inspector's job was to tick a box: yes, or no. Under SIRE 2.0, that binary is gone entirely — replaced by a four-tier performance rating built from observation, evidence, and context. For crews and superintendents used to the old system, this is often the single hardest adjustment to make, because it changes not just what's being asked, but what counts as a good answer.
What You'll Learn in This Lesson
- Why SIRE 2.0 replaced pass/fail with a four-tier rating system
- What separates each of the four performance ratings, with real examples
- The step-by-step process an inspector actually follows to reach a rating
- Why "evidence vs opinion" is the most common point of dispute in an inspection
- Role-specific best practices for crew, Master, Chief Officer, and Chief Engineer
Why SIRE 2.0 No Longer Uses Simple Pass/Fail
VIQ7's yes/no format made for a fast report, but a shallow one. A "yes" told a charterer that a box had been ticked; it told them almost nothing about how close that box came to not being ticked, or how the crew would perform under slightly different conditions. SIRE 2.0's four-tier system exists to close that gap — trading a small amount of simplicity for a much more useful signal about actual operational risk.
The result is a report that reads less like a checklist and more like a graded assessment, where the difference between "acceptable" and "excellent" is recorded, not collapsed into the same tick.
Observation-Based Assessment: The Four Building Blocks
Every rating an inspector assigns rests on four things coming together:
- Objective evidence — what the inspector directly saw, heard, or was shown, not what they were told in the abstract.
- Observation — the specific, recorded instance the rating is based on, tied to a place, a task, and a moment.
- Context — the conditions surrounding that observation: was the crew mid-drill, under a real operational load, newly assigned to the task?
- Risk — what the gap between actual and ideal performance would mean in a real incident scenario, not just whether a rule was technically followed.
An inspector who can't tie a rating back to objective evidence, a specific observation, its context, and its risk implication hasn't finished their assessment — they've formed an impression. SIRE 2.0 is built to prevent impressions from becoming ratings.
The Four Performance Ratings, With Real Examples
Every Human Factors, Hardware, and Procedures question in a CVIQ is scored on the same four-tier scale:
Exceeds Expectations
Performance that goes beyond the minimum standard in a way that visibly reduces risk. Example: a junior engineer, asked to explain a fuel-system isolation procedure, not only describes the correct steps but proactively flags a labelling inconsistency on the panel that could confuse someone under pressure.
As Expected
The standard result: the task is performed correctly, the reasoning is sound, and nothing about the observation raises concern. Example: an AB completes a mooring operation following the documented procedure, communicates clearly with the bridge, and can explain why each step matters.
Largely as Expected
Minor weaknesses that don't compromise safety but are worth noting. Example: an officer performs a fire drill correctly but hesitates when asked to explain the reasoning behind one specific step, recovering the correct answer only after a prompt.
Not as Expected
A meaningful gap between actual and required performance, tied to specific evidence. Example: a crew member cannot explain why a permit-to-work step exists, performs it purely by rote, and cannot adapt when the inspector varies the scenario slightly.
How Inspectors Actually Decide
Behind every rating is a repeatable sequence, not a snap judgement:
A CVIQ question is asked, the inspector observes the task being performed, interviews the crew member about their reasoning, gathers objective evidence (photos, documents, direct demonstration), identifies any relevant Performance Influencing Factor, arrives at one of the four ratings, and records it as a formal observation in the report. Skipping a step in this sequence is exactly how disputes over a rating start.
Download the Observation Rating Matrix (PDF)
A one-page reference mapping all four rating tiers to real evidence examples across Hardware, Procedures, and Human Factors.
We'll also send you one maritime compliance update per week. Unsubscribe anytime.
Evidence vs Opinion: Where Most Disputes Start
This is the section IMT spends the most time on with superintendents, because it's where the most avoidable disputes originate. An inspector's rating should always be traceable to something objective — a photo, a document, a demonstrated action, a direct answer to a direct question. An operator's disagreement with a rating is only productive if it's argued the same way: with evidence, not with a general assertion that "our crew knows this."
The most common failure mode on the operator side is treating a rating as a matter of opinion to be negotiated, rather than a claim about evidence to be checked. If a "Not as Expected" rating is wrong, the way to challenge it is to identify what evidence the inspector missed or misread — not to argue that the crew is generally competent.
Common Mistakes
- Responding to a low rating with a character defence of the crew member instead of a specific evidence challenge.
- Assuming "As Expected" is a weak result — it is, in fact, the correct and desired outcome for the overwhelming majority of observations.
- Failing to document context (fatigue, first time performing the task, equipment recently changed) that would have supported a fairer rating.
- Preparing crew to give "correct" answers without preparing them to demonstrate the reasoning an inspector is actually testing for.
Best Practices, by Role
- Crew and ratings — be ready to explain the "why" behind a routine task, not just perform it; small hesitations under questioning are a common source of "Largely as Expected" ratings that are otherwise avoidable.
- Master — ensure context that could affect a rating (recent crew changes, equipment updates, unusual operational conditions) is proactively shared with the inspector, not left for them to discover.
- Chief Officer — treat deck-side procedure documentation as something the crew actually references, not paperwork maintained solely for inspection.
- Chief Engineer — keep machinery-space labelling and documentation current enough that a junior engineer can reference it under questioning without hesitation.
Need a Second Opinion on a Rating?
Talk to an IMT compliance advisor about reviewing a recent SIRE 2.0 observation or preparing an evidence-based response.
Contact an ExpertSummary
SIRE 2.0's four-tier rating system replaced a fast but shallow pass/fail check with a slower, evidence-based process built on objective evidence, direct observation, context, and risk. Every rating traces back through a repeatable sequence — question, observe, interview, evidence, PIF, rating, observation — and the ratings themselves range from Exceeds Expectations down to Not as Expected. The single most useful habit for any operator is treating every rating as a claim about evidence, and responding to it the same way.
Get Your Crew SIRE 2.0 Ready
IMT's SIRE 2.0 Readiness course walks your officers and crew through real CVIQ scenarios, human factor interview technique, and hardware documentation practice.
View the Course Talk to an AdvisorDownload Center
Everything from this article, in a format you can print, share, or file.
Frequently Asked Questions
Which rating indicates acceptable performance with minor weaknesses?
"Largely as Expected" — the task is performed acceptably, but with small gaps that don't compromise safety, worth noting rather than escalating.
Is "As Expected" a bad result?
No — it's the standard, correct outcome for most observations and should be read as a positive result, not a missed opportunity for something higher.
What's the difference between an observation and a rating?
The rating is the score (one of the four tiers); the observation is the recorded evidence and context behind that score.
How should an operator dispute a rating it disagrees with?
By identifying specific evidence the inspector may have missed or misread — not by asserting general confidence in the crew's competence.
Does context ever change a rating after the fact?
Context should ideally be surfaced during the inspection itself; operators who proactively share relevant context tend to see fewer disputed ratings later.
Test Your Knowledge
5 questions · pass with 4/5 to unlock your Scoring & Ratings certificate.
Congratulations!
Badge Earned — Scoring & Ratings