Skip to main content
03 · AIRE Analyst™ANA · Rigor/Execution

AIRE Analyst™

Technology expression: Eval Analyst

"The one who makes the numbers defensible."

Your result has not changed — Eval Analyst is the Technology expression of your AIRE Analyst™.

TL;DR — You are the person who asks what the number actually measures.

When someone says the new prompt is better, you want the sample, the baseline, and whether the difference would survive a rerun.

You have seen enough demonstrations chosen from a good afternoon of outputs to know that a screenshot is not evidence.

Full Profile

The complete Eval Analyst analysis

Core Drive

You are driven to put a number on the eval or the offline metric that will still be true after the model review leaves the shared page. On an eval harness, prompt comparison, or online guardrail claim that means every figure traces to a defined metric, a comparison set, and an assumption you can defend. You are the colleague who asks what the number actually measures before you let it onto the eval sheet, the launch memo, the guardrail ticket, or the model-review pack. You measure success in claims caught in the assumption log, not in being the most confident person in the model review.

How You Work

You work by treating the model as a second analyst that has to survive last quarter's closed eval files. You load recent offline metrics, slice failures, contamination notes, and the metric definitions the pack already treats as settled, then ask which lines move when the comparison window changes or the set is one prompt version versus the baseline harness. You immediately rerun the same query against a hold-out slice, a prior model version, or a second traffic window so a hallucinated lift or guardrail rate dies before it hits the shared page or the model review. Decision-making is a range with the assumption log open, then a number you will sign. Communication is one page a model review or launch huddle can use, not a workbook. You iterate by changing one variable (metric definition, contamination filter, slice, time window, online vs offline pairing) and watching whether the range still closes. You do not paste unpublished training corpora or non-public customer transcripts into an unapproved tool; eval claims stay on approved harnesses and disclosed pack data inside the workflow. Pack-backed figures only land on the shared page.

Your Strengths

Your name is on the number, so you refuse to carry what you cannot trace. You break a claim that the new prompt is better — or that the guardrail should ship at the model's mid-point — into what the harness logs versus what the offline metric and online slice actually show. You flag when a demo sample or a pulled composite violates last quarter's comparison set. You keep the assumption log a reviewer, PM, or on-call lead can follow. You translate a model output into a one-pager that belongs in the model review, not in a slide. You are interested in anything that speeds eval lookup, slice checks, or gap detection, and you kill anything that prints an eval figure with no definition behind it.

Blind Spots

The same rigor can freeze a model review or a launch huddle while the ship window is still open. You may demand another week of offline metrics after the train already needs a guardrail decision tonight. You can discount a practitioner's warning that the slice will not behave like the comps because the spreadsheet is green, and lose the call to someone willing to name what the grid cannot see.

Under Pressure

When the model-review pack is this afternoon, a launch is tomorrow, or a guardrail stack is still open on meeting day, you open more models rather than fewer. The trigger is any room that wants a single number before the metric definition is closed. In those moments you may withhold the finding until the range is prettier, and the window to actually set the eval gate or the online guardrail closes.

On a Team

Engineers, PMs, and reviewers come to you because the eval or guardrail claim feels real instead of dashboard-smooth. They trust your ranges. They sometimes wish you would say the model is fine without another tab. You fill the role of the person whose number survives a harness question, a slice question, or a model-review question.

AI Connection

You adopt AI the moment it traces an eval line to a metric definition and a comparison you can check against a prior window. You resist tools that emit an offline or online figure with no assumption log. Once a model matches your historical pattern on one eval pack, you lock that prompt into the harness template and move to the next slice. You stay inside the approved-tool list and keep unpublished corpora and non-public transcripts out of unapproved models.

Famous Parallels

The eval analysts who asked whether a dashboard tile measured offline accuracy or online guardrail rate, the prompt analysts who rebuilt a lift claim from the actual harness set instead of from the demo tile, and the quality analysts who rebuilt a slice recommendation from recent contamination notes instead of from the model's mid-point.

One-Liner

"Option A I can trace to a metric definition and a comparison set. Option B still has an open definition. I'm not signing B."

Your Strengths

  • ✓You notice wrong output faster than the people around you, because you are checking the logic rather than the presentation.
  • ✓You build methods other people can re-run and get the same answer from.
  • ✓You are the person whose number is trusted when the number actually matters.
  • ✓You naturally put a value on a tool — hours saved, errors avoided — instead of arguing about it in the abstract.

Your Blind Spots

  • ◐You tend to adopt only where verification is easy, and skip the areas where a tool could save the most time.
  • ◐Your pace is often slower than the decision needs, even when the extra precision changes nothing.
  • ◐The important finding can get buried in the detail supporting it.
  • ◐Small uncertainties and serious ones get treated with the same weight.

Illustrative AIRE Radar

Rigor84
Execution74
Awareness59
Initiative53

Illustrative only — Rigor 84, Execution 74, Awareness 59, Initiative 53. Take the assessment to see your actual A/I/R/E scores.

For Employers

Someone has to say whether the output is actually better, or only faster. Measurement and business case — attach them to any initiative that has an ROI claim.

Your 30-Day Action

Take one live eval or offline metric already in motion (an eval harness pack, a dashboard tile people treat as the ship signal, an online guardrail line, or a prompt-lift recommendation). Ask what the number actually measures versus what people think it measures. Log the metric definition, the comparison window, and the adjustments; keep unpublished corpora inside the approved workflow and put only pack-backed figures on the shared page. Verifiable check: a one-page assumption log attached to that eval pack or launch memo is used in one model review or launch huddle within 30 days, with each line tracing the number to a named metric definition.

Explore the other Technology expressions

Haven't taken AIRE yet?

41 questions · ~8 minutes · free to take · $7.99 to unlock your full report.

Start the assessment →