Skip to article
THE AI HIRING FIELD GUIDE / 01

The answer looks right. Can your candidate explain why?

A practical guide to hiring people who can build with AI, spot what went wrong, and take responsibility for the result. Six skills. Better questions. Clearer evidence.

Build a better AI skills assessment
THE HIRING PROBLEMVERIFY THE OUTPUT
THE SAME BRIEF. TWO FINISHED ANSWERS.
“Your request is eligible for a refund.”
Both answers look convincing. Only one is verified.
Stock portrait representing candidate ACANDIDATE A“The model
said so.”

Accepts the answer.
Moves on.

Understanding still untested
Stock portrait representing candidate BCANDIDATE B“Let’s check
the source.”

Finds an old policy.
Corrects the answer.

Judgment becomes visible

Illustrative scenario. Stock portraits, not real applicants.

A practical guide for recruiters, talent leaders & engineering managers Get the interview kit
In this guide

Two candidates. The same brief. Two convincing answers. Your scorecard says they performed equally well. Your hiring manager still needs to know whom to trust with the work.

That is the uncomfortable part of assessing AI skills: the finished output can hide the decisions that produced it. A candidate might understand every choice. They might have accepted the first plausible response. The artifact alone may not tell you which.

For a recruiter, this creates an awkward handoff. You can report a score, but the engineering manager asks, “Could they do this when the example changes?” For the engineering manager, it creates another interview to reconstruct what the assessment should have revealed.

Here is the change we recommend: give candidates a useful task, then introduce a reason to question the answer. Observe how they frame the problem, build a solution, verify it, and explain a decision. This guide shows how to do that across six AI skill areas.

01 / CHOOSEWork from the role.Pick a task the person might actually own.
02 / OBSERVEChange one condition.Make judgment visible when the easy answer breaks.
03 / DISCUSSAsk for the evidence.Connect the result to a test, a source, or an explanation.

What AI fluency actually means

For hiring purposes, AI fluency is the ability to frame a task, direct AI-assisted work, evaluate the result, and improve it when the evidence changes. It combines tool use with responsibility for the outcome.

It is different from listing tools on a resume. A person can be familiar with a coding assistant and still struggle to notice that its change breaks another part of the application. Another person may use a less familiar tool but work methodically: inspect the context, make a focused change, test it, and explain what remains uncertain.

The assessment should help you see those behaviors. It should also match the role. You probably do not need a frontend engineer to design a retrieval pipeline from scratch. You may need that engineer to recognize when a generated component mishandles an API response.

A useful opening question for the hiring team is: “What AI-assisted work will this person need to own in their first few months?” Write the answer in ordinary language before choosing a test format.

A working definition for your scorecard

“Can use AI to produce useful work, check the result against evidence, and explain the important decisions.” Add the technical depth required by the role.

Make the reasoning observable

Start with one bounded task. Give the candidate a clear outcome, the permitted tools, and enough context to begin. Then plan what evidence you want to collect at each stage.

FIG. 01 / THE EVIDENCE LOOPOne task. Four things worth observing.
FRAME

What does good look like?

Define the user, the constraint, and a testable outcome before touching the model.

Evidence: a clear task and acceptance criteria

The diagram is a review framework, not a claim that every candidate must follow one exact sequence. Good work can involve returning to the brief, testing early, or discovering that the first approach was unnecessary.

Ask candidates to explain decisions they can discuss directly. You do not need to infer their private thinking from pauses, typing speed, or a recording. The useful evidence is observable: a test they ran, a source they checked, a change they made, or a reason they can defend.

This is related to a useful principle from AI evaluation: examine both the outcome and the path taken to reach it. Anthropic's discussion of agent evaluations distinguishes final outcomes from the interaction trace used to understand behavior. A hiring assessment needs its own role-specific criteria, but the distinction is useful here. Read the evaluation background.

The most useful moment in an AI assessment may be the moment the first answer stops working.

Six skills, six ways to test them

These categories give your team a vocabulary for the work. They overlap. Select the ones that matter for the role instead of turning the entire list into a marathon assessment.

Six AI skills grouped into three layers: prompt engineering and AI-assisted coding to direct the work; agents and RAG to connect the system; harnesses and control loops to make work repeatable.
FIG. 02 / A working map of AI skills. These layers overlap; choose the relevant depth for the role.

1. Prompt engineering: can they define a useful result?

A prompt-engineering task should reveal whether the candidate can turn a goal into instructions, context, constraints, and a result you can evaluate. “Write a clever prompt” is too vague to give reviewers much agreement about success.

Try this task: Give the candidate several sample invoices and a target JSON structure. Ask them to extract the required fields. Include one missing due date and one document with conflicting totals.

Look for how they define missing information, distinguish a document value from a guess, and check the output format. A longer prompt is not automatically a stronger prompt. Ask what each important instruction is intended to change.

Change one thing: Introduce an invoice layout they have not seen. Does the candidate diagnose the failure or simply add more instructions without checking which one helps?

Ask: “Which example would you use to show that your revision actually improved the prompt?”

2. Agent development: can they control actions as well as answers?

A tool-using agent can affect a system, so an assessment should explore the conditions under which an action is allowed and how failures are handled. The candidate needs to reason about the workflow around the model.

Try this task: Build a small support workflow that looks up an order and proposes the next action. Keep it in a sandbox with simulated tools. Include a lookup failure and an order that needs human review.

Look for input validation, a clear relationship between evidence and action, and a response when the required information is unavailable. An impressive number of agents or tools is not a scoring criterion.

Change one thing: A tool times out after a request may already have succeeded. Ask what the system should do before trying again.

Ask: “Where would you use a fixed workflow rather than let the model decide the next step?” Anthropic's distinction between predefined workflows and model-directed agents is useful technical background for this question. Explore the architecture distinction.

3. RAG pipelines: can they connect an answer to the right evidence?

Retrieval-augmented generation, or RAG, combines retrieved information with generated responses. The original RAG research describes combining parametric generation with retrieved non-parametric memory. In a hiring task, the practical question is whether the candidate can get the relevant material into the answer and check that it supports the claim. Read the original research.

Try this task: Provide a small policy collection and ask the candidate to answer employee questions with supporting references. Include an old policy, its replacement, and a question the collection cannot answer.

Look for source selection, treatment of outdated information, and a sensible response when the evidence is insufficient. A citation that looks tidy is not enough; it needs to support the answer.

Change one thing: Add a similar-looking document that applies to a different department.

Ask: “Did the wrong answer come from retrieval, the supplied context, or generation? How would you distinguish them?”

4. Harness engineering: can the work resume without losing its footing?

Here, a harness means the surrounding environment that organizes an agent's work: tools, useful context, state, progress records, and verification. Teams use the term differently, so spell out the scope in the assessment brief.

Try this task: Give a candidate an unfinished coding-agent session and a small project. Ask them to design how another session would resume safely: what should be recorded, what should be inspected, and how completion should be checked.

Look for a distinction between “the agent said it finished” and evidence that the change exists and works. A progress note is useful when it points to verifiable state.

Change one thing: The progress record says a task is complete, but the relevant check fails.

Ask: “Which source of state do you trust, and what would you verify before the agent continues?” Anthropic's work on long-running harnesses offers background on initialization, progress records, and incremental work across sessions. Read the harness discussion.

5. Loop engineering: can they make another attempt worthwhile?

For this guide, loop engineering means shaping the repeated cycle of action, observation, evaluation, and correction. It is a working assessment category, not a universal job-title definition.

Try this task: Show an agent repeatedly rewriting a function while the same test keeps failing. Ask the candidate to improve the feedback it receives and define when it should stop.

Look for a concrete diagnosis, useful feedback, a bounded retry strategy, and a reason to escalate. “Try again” does not tell the agent what it should learn from the previous attempt.

Change one thing: The latest revision passes the original test but fails a regression test.

Ask: “What evidence justifies another attempt, and what evidence means a person should intervene?” The evaluator-optimizer pattern is one technical reference for thinking about feedback and iteration. See the pattern.

6. AI-assisted coding: can they own what the assistant writes?

A role may involve Cursor, Claude Code, Codex, Antigravity, or another coding environment. Establish the tools candidates will actually have before you send the task. Tool familiarity and engineering judgment should be distinguishable in the review.

Try this task: Ask the candidate to add validation to a small existing API using the permitted assistant. Provide a clear requirement and a few tests. Review the change, including anything unrelated that the assistant modified.

Look for project inspection, specific instructions, review of the resulting diff, and meaningful verification. The candidate should be able to explain the code even if they did not type every line.

Change one thing: Introduce a duplicate request or an empty payload that the happy-path test did not cover.

Ask: “What did you accept from the assistant, what did you change, and what convinced you the result was ready?”

A worked example: the refund answer

Imagine you are hiring an AI application engineer. Your team wants a policy assistant that can answer a straightforward customer question without inventing a policy.

The candidate receives a small document set and a request: “I bought this 21 days ago. Am I eligible for a refund?” One older document allows returns within 30 days. The current policy allows 14 days. For this exercise, the current policy applies to the purchase and there are no additional exceptions.

Both initial outputs say the customer is eligible. That is the starting point in the animation below, not a correct final answer.

ONE TASK. A REVEALING MOMENT.
Convincing is not the same as checked.

The assistant uses an archived 30-day policy. The answer sounds reasonable, but the source is wrong.

14-second loop · Illustrative scenario, stock portrait. Full worked example below.

A surface-level review might reward the clear wording and a reference to a policy. A deeper assessment asks the candidate to check which policy applies.

The candidate with stronger evidence notices the date conflict, uses the current policy, and changes the answer. They then add a check that would catch the same failure later. Their final response explains that the request is outside the 14-day window under the stated assumptions.

What should you record? The document they selected, the reason they selected it, the change in the answer, and the test or check they proposed. You now have something specific to discuss with the hiring manager.

What should you avoid inferring? This one task does not establish that the person can operate every retrieval system. It demonstrates a bounded piece of judgment. The next round can explore the technical depth the role still needs.

A group of colleagues reviewing work together; illustrative stock photograph
The goal is a better conversation between recruiting and engineering. Illustrative stock photography, not a customer testimonial.

Match the assessment to the role

A recruiter should not need to become an expert in all six categories to run a useful intake meeting. Ask the hiring manager to name one representative task, one failure that matters, and the evidence they would need to trust the result.

Use those answers to choose a small set of skills. If every requirement appears equally important, the brief probably needs another conversation before it needs another question.

BUILD THE ASSESSMENT AROUND THE ROLEChoose the work before you choose the questions.

Repair a policy-answering assistant

Assess
Prompt engineering + RAG + code review
Change one thing
Introduce an outdated policy and one unanswered question.
Look for
Source selection, grounded answers, a regression test.

Illustrative blueprints. Set difficulty, duration, and tool access for your actual role.

For an AI application engineer, retrieval quality and verification may be central. For an agent engineer, tool actions, state, and recovery may deserve more attention. For a software engineer, supervising a coding assistant can be evaluated alongside programming fundamentals.

Seniority changes the expected depth. A junior candidate may identify a bug and describe a sound check. A senior candidate may need to explain the operational consequences and the limits of the approach. Use the same role-level expectations for everyone taking that assessment.

Use a scorecard you can explain

An overall score can organize a shortlist. The hiring discussion still needs the evidence beneath it. Before candidates begin, agree on what demonstrated ability looks like for each important dimension.

The worksheet below uses four dimensions: framing, implementation, verification, and explanation. It deliberately avoids a universal pass mark. Your team needs to calibrate the criteria for its role rather than inherit a number from a blog post.

TRY THE REVIEW WORKSHEET

What did you actually observe?

0/4 reviewed

This is an example rubric, not a validated hiring cutoff. Select an observation for each dimension.

Problem framing

Starts without clarifying the outcome.
Stronger evidence: Defines the outcome and important constraints.

Implementation

Produces output but cannot establish correctness.
Stronger evidence: Produces a working solution for the stated requirements.

Verification

Accepts a plausible result without checking.
Stronger evidence: Tests an edge case and connects evidence to a correction.

Explanation

Repeats what the tool produced.
Stronger evidence: Explains a choice and adapts it when requirements change.

Write observations that another reviewer can inspect. “Strong at AI” is difficult to challenge or build on. “Identified the outdated document and added a regression case” tells the next interviewer what was demonstrated.

If reviewers disagree, locate the disagreement. Are they weighing the role's requirements differently? Did one reviewer see evidence the other missed? Is a criterion too vague? A calibration conversation is more useful than averaging two unexplained judgments.

Treat missing evidence as a question to resolve. It may justify a focused follow-up rather than an immediate negative conclusion.

Run a focused assessment and follow-up

Here is an illustrative 45-minute format, suitable as a starting point for a bounded task. Pilot it with your team before using it with applicants. More complex work needs a different scope, not a promise that everyone should finish at this pace.

  1. First 5 minutes: frame the problem. Let the candidate inspect the brief, clarify the desired outcome, and understand tool access.
  2. Next 20 minutes: build or repair. Give them a task small enough to produce reviewable work.
  3. Next 10 minutes: change one condition. Introduce a relevant edge case, source conflict, or failed check and observe their response.
  4. Final 10 minutes: discuss. Ask about one important decision, one unresolved risk, and what they would verify next.

Tell candidates about the format in advance. Explain that the follow-up explores their work and that acknowledging a limitation is acceptable. Arrange appropriate accommodations and a way to report environment problems through your hiring process.

Three questions to take into the next interview

“What convinced you this was correct?” Ask for the source, test, or observation, not a general assurance.

“What changed your approach?” Look for a connection between new evidence and a revised decision.

“What would you check before someone relied on this?” Explore whether the candidate can prioritize the consequences that matter for the role.

You do not need a trick question. A small, relevant change to the task can give you plenty to discuss.

Take the questions into your next intake meeting.

Download the role brief, sample task, review rubric, and follow-up prompts as an editable Markdown kit.

Get the interview kit ↗

Where LunaPrompts fits

LunaPrompts brings practical AI skills assessment into a technical hiring workflow. Teams can combine relevant task formats, review candidate evidence, and use follow-up explanations to prepare a more focused interview.

The useful product conversation starts with your role. Bring the work the candidate will own, the tools they will use, and the questions your current process leaves unanswered. From there, define the assessment and the evidence your reviewers need.

AI-assisted work and integrity review also need consistent rules. If a coding assistant is permitted, using it is part of the task. If independent work is required, say so before the candidate starts. A flag or similarity indicator needs context; it is not a substitute for investigating the candidate's work and explanation.

Explore LunaPrompts assessment formats and candidate analytics, or bring a role to a pilot discussion.

Questions hiring teams ask

Can we assess AI skills without hiring an AI specialist?

Yes, if the assessment matches the role. A general software engineer may need to supervise AI-generated code and verify a change. That does not mean they need to build a specialized agent harness. Choose the depth from the responsibilities.

Is prompt engineering enough to establish AI fluency?

A prompt task can provide useful evidence about instructions and evaluation. Broader AI work may also require implementation, retrieval, tool use, debugging, and verification. Test the combination the job requires.

Should we let candidates use AI during a technical assessment?

Use permitted AI when you want to evaluate AI-assisted work. Use a clearly identified independent task when you need separate evidence of foundational ability. Explain the rules, tools, and review criteria for each stage.

Does a candidate need to be a polished speaker?

The relevant signal is whether they can explain the technical decision well enough for the role. Ask focused follow-ups and consider the work alongside the explanation. A confident delivery alone does not establish correctness.

Does this framework predict who will be a good hire?

It is a practical way to gather relevant evidence, not a validated predictor or an automatic hiring decision. Calibrate it for the role, review how it performs, and combine the result with the other evidence your process needs.

One question to leave with

Before sending your next AI assessment, ask: “Where will the candidate have to notice that the first answer is not good enough?” If the task never creates that moment, you may still be measuring the output more than the person.

Sources & editorial notes

Technical background informs the definitions in this guide. The hiring tasks, four-part review framework, and 45-minute format are LunaPrompts editorial examples, not conclusions from these sources. Scenarios and candidate portraits are illustrative. Written October 11, 2026.

  1. Anthropic: Building effective agents. Workflow, agent, and feedback-loop architecture.
  2. Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Research background for RAG.
  3. Anthropic: Effective harnesses for long-running agents. Progress, state, and verification across sessions.
  4. Anthropic: Demystifying evals for AI agents. Outcomes, traces, and evaluation design.
KEEP EXPLORINGAll 19 field guides
PUT THE GUIDE TO WORK

Bring the role.
We’ll help you find the signal.

Explore a hiring workflow built around the work your next hire will actually do.

Explore your hiring workflow
Colleagues discussing work around a table; illustrative stock photography