← Back

22 Output Evaluation & Validation Practice Questions & Answers

Every Output Evaluation & Validation practice question from the Claude Certified Associate – Foundations Practice Test, with the correct answer and a short explanation.

Start practice test
  1. 1. Priya, a project manager, asks Claude to summarize a vendor contract. The summary is fluent, confidently worded, and cites Section 8.4(b) and a 45-day termination notice. Priya is about to forward it to Legal. What is the MOST appropriate next action?

    • A.Ask Claude to rewrite the summary in more formal legal language before forwarding
    • B.Open the contract and check Section 8.4(b) and the 45-day figure against the actual text before forwardingAnswer
    • C.Ask Claude how confident it is in the citation and forward it if Claude reports high confidence
    • D.Forward it as is, since the summary is detailed and states the section number confidently

    A cited section number plus a precise figure is exactly the kind of specific-looking detail most prone to fabrication, and a model's expressed confidence is decoupled from its accuracy, so self-rated confidence is not a validation technique. Reformatting changes style, not facts; only comparison against the authoritative source document verifies the claim.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 3 (apply fact-checking and validation techniques); Anthropic sample item rationaleReport a problem with this question

  2. 2. Amara, an operations lead, uploads a 300-page policy manual and asks Claude a detailed compliance question. Which approach BEST grounds the answer in the manual?

    • A.Ask the question in several different wordings and keep whichever answer is longest
    • B.Have Claude first extract the word-for-word passages relevant to the question, then reason only from those extracted quotesAnswer
    • C.Paste only the manual's table of contents to save space, then ask the question
    • D.Ask the question directly and accept the answer if it reads as well organized and specific

    For long documents Anthropic recommends direct-quote grounding: have the model pull verbatim passages first and then reason only from those quotes, which makes every downstream claim traceable to text you can check. Length, organization and specificity of prose are not accuracy signals.

    Source: Anthropic docs, Reduce hallucinations: direct quote grounding for long documentsReport a problem with this question

  3. 3. A consultant uploads three client reports and asks a question that none of them actually covers. Claude produces a smooth, plausible answer anyway. What is the BEST way to prevent this on future queries?

    • A.Explicitly instruct Claude to answer only from the uploaded reports and to state when they do not contain the answerAnswer
    • B.Ask Claude to add a confidence percentage to each paragraph of its answer
    • C.Tell Claude to be more accurate and to avoid making things up
    • D.Switch to the most capable model available and ask the same question again

    Two named techniques combine here: external knowledge restriction (use only the supplied documents) and explicitly allowing the model to say it does not know, which removes the pressure to fill a gap with invention. A vague plea for accuracy, self-scored confidence, and escalating model tier do not address a grounding defect.

    Source: Anthropic docs, Reduce hallucinations: allow I don't know + external knowledge restrictionReport a problem with this question

  4. 4. A marketing associate has a Claude-drafted white paper containing twelve statistics, all attributed to an industry study that was supplied as a PDF. Which PAIR of steps is the MOST reliable verification?

    • A.Require a supporting verbatim quote for each statistic, then remove any statistic the source cannot supportAnswer
    • B.Ask Claude to double-check its own numbers, then publish if it says they are fine
    • C.Ask Claude to restate the statistics as approximations, then publish without checking the PDF
    • D.Spot-check the first statistic, then assume the rest follow the same pattern

    Verify-with-citations means every factual claim must be tied to a quote from the source, and any claim that cannot be supported is retracted rather than softened. Self-checking inside the same thread, sampling one item, and hedging the wording all leave unsupported numbers in the published document.

    Source: Anthropic docs, Reduce hallucinations: verify with citations (retract unsupported claims)Report a problem with this question

  5. 5. An analyst needs one number from Claude for a board slide and has no primary source at hand. Which check BEST signals that the number may be fabricated?

    • A.Run the same prompt several times in fresh conversations and see whether the number stays consistentAnswer
    • B.Check whether the answer is written in complete, professional sentences
    • C.Ask Claude to repeat the number in the same conversation to see if it agrees with itself
    • D.Ask Claude whether it is certain, and treat a yes as confirmation

    Best-of-N verification repeats the same prompt independently and compares results, because a fabricated specific tends to vary across runs while a genuinely grounded one does not. Repeating inside the same conversation only echoes the earlier answer, and neither prose quality nor self-declared certainty tracks accuracy.

    Source: Anthropic docs, Reduce hallucinations: Best-of-N verificationReport a problem with this question

  6. 6. A finance coordinator receives a Claude-built cost model whose final total looks wrong, though each assumption is stated clearly. What is the BEST next step to find the defect?

    • A.Ask Claude to try again from scratch without saying what looked wrong
    • B.Accept the total, since the assumptions are documented and the write-up is clear
    • C.Ask Claude to lay out its calculation step by step before restating the total, so the faulty step is visibleAnswer
    • D.Ask Claude to present the same model as a polished executive summary

    Chain-of-thought verification asks the model to expose its reasoning steps before the conclusion, which surfaces the specific arithmetic or assumption error instead of hiding it behind a clean result. Reformatting is cosmetic, documented assumptions do not guarantee correct arithmetic, and a blind regeneration does not name the defect.

    Source: Anthropic docs, Reduce hallucinations: chain-of-thought verificationReport a problem with this question

  7. 7. An HR associate asks Claude to answer employee questions from the company benefits handbook. Several answers mix in general industry practice that is not in the handbook. What is the MOST appropriate fix?

    • A.Instruct Claude to use only the handbook as its source and to say when the handbook is silent on a questionAnswer
    • B.Ask Claude to write the answers in a warmer, more employee-friendly tone
    • C.Stop using Claude for benefits questions entirely
    • D.Add a disclaimer at the bottom of each answer saying details may vary

    The defect is a missing grounding constraint, so the fix is external knowledge restriction plus permission to report gaps, which keeps answers traceable to the authoritative handbook. A disclaimer and a tone change leave the wrong content in place, and abandoning the task discards a workflow that a single safeguard makes safe.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 3; Anthropic docs on external knowledge restrictionReport a problem with this question

  8. 8. A program manager asked for a project brief covering scope, timeline, budget, risks and stakeholders. Claude returns a polished brief with no risks section, and everything present is factually correct. How should she evaluate this output?

    • A.As acceptable, because everything stated in it is factually accurate
    • B.As an accuracy failure requiring the whole brief to be regenerated from scratch
    • C.As acceptable, because the writing quality shows the model understood the task
    • D.As a completeness failure: the output is accurate but omits a required element, so it does not meet the task requirementsAnswer

    Accuracy and completeness are distinct evaluation axes: an output can be entirely correct and still fail because a requested element is missing. Evaluation is against the original task requirements, not against how polished the prose reads, and the targeted fix is to request the missing section rather than discard correct work.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 1 (evaluate outputs for accuracy and completeness)Report a problem with this question

  9. 9. An operations lead asked why customer churn rose across all four regions last quarter. Claude returns a detailed, well-argued explanation about the Northeast region only. What is the BEST next action?

    • A.Accept it, since the Northeast analysis is thorough and well supported
    • B.Re-run the original prompt unchanged and hope for broader coverage
    • C.Reply naming the defect: the answer covered one of four regions, and ask for the same analysis for the remaining threeAnswer
    • D.Ask Claude to make the existing answer longer and more formal

    The output silently narrowed the scope of the question, answering an easier nearby question instead of the one asked, which is a scope-drift defect rather than a quality-of-writing problem. Effective iteration names the specific defect in a follow-up turn; blind re-rolling and lengthening address neither the omission nor its cause.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 7 objective 1 (diagnose and resolve poor outputs)Report a problem with this question

  10. 10. A knowledge worker must judge whether a Claude-generated policy summary is good enough to circulate. What should she do FIRST?

    • A.Ask Claude to grade its own summary against a scale of one to ten
    • B.Circulate it to colleagues and let their reactions decide whether it is good enough
    • C.Read the output first and form an overall impression of whether it feels solid
    • D.Write down the acceptance criteria the summary must meet, then read the output against that listAnswer

    Criteria defined before reading prevent fluency and confident tone from standing in for substance, because an impression formed while reading is already anchored by the writing quality. Self-grading is not a validation technique, and circulating unverified content shifts the review burden after the risk has been taken.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 5 (compare outputs against explicit criteria)Report a problem with this question

  11. 11. A communications specialist has two Claude drafts of the same customer notice and must pick one. What is the MOST appropriate way to choose?

    • A.Pick the longer draft, since it likely covers more ground
    • B.Score both against the same explicit rubric of accuracy, completeness, tone fit and actionabilityAnswer
    • C.Pick the one that reads more smoothly on a first pass
    • D.Ask Claude which of its two drafts is better and use that one

    Comparing candidates requires a shared, pre-stated rubric so the decision rests on defined dimensions rather than on which draft happens to read better. Smoothness and length are not quality signals, and asking the model to rank its own outputs reintroduces the same unreliable self-assessment.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 5 (compare outputs for the intended audience)Report a problem with this question

  12. 12. A support manager fixed a prompt after Claude mishandled one tricky refund ticket. What is the BEST way to confirm the revised prompt is actually better?

    • A.Re-run the revised prompt against a small set of real tickets that previously failed, plus typical ones, and compare all resultsAnswer
    • B.Assume the fix worked, since the new instruction directly addresses what went wrong
    • C.Re-run only the ticket that prompted the change and ship the prompt if it now passes
    • D.Add several more instructions to the prompt so it covers every case he can imagine

    A small regression set of representative real failures plus normal cases catches fixes that solve one case while breaking others, which a single-case retest cannot detect. Stacking extra instructions onto a prompt without testing tends to introduce new conflicts rather than confirm an improvement.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 7 objective 2 (adjust approach based on results)Report a problem with this question

  13. 13. A research associate receives a competitor overview that is accurate but written at a depth far beyond what the executive audience needs. What is the MOST effective follow-up?

    • A.Ask Claude to make the document sound more professional
    • B.Manually delete paragraphs until the length looks right, without checking what was lost
    • C.State the specific defect, that the detail level is wrong for an executive reader, and specify the length and level of detail requiredAnswer
    • D.Re-run the original prompt and hope for a shorter result

    Effective iteration diagnoses the specific defect and states the corrected requirement, here audience-appropriate depth and length, so the next output targets the real gap. Blind regeneration, unchecked manual cutting that may drop caveats, and a vague professionalism request all fail to specify what needs to change.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 5 (edit, adapt and refine for the intended audience)Report a problem with this question

  14. 14. An associate notices Claude's output on a routine task is generic and misses the company's context. What is the MOST appropriate first fix?

    • A.Tell Claude the previous answer was bad and ask it to try harder
    • B.Switch to the most capable and expensive model tier available
    • C.Ask the same question repeatedly until a better answer appears
    • D.Supply the missing background, the intended reader and explicit success criteria in the requestAnswer

    Generic output is a symptom of missing context, audience and success criteria, so the root-cause fix is to supply them rather than to escalate model tier, which costs more without addressing the defect. Repetition and vague criticism give the model no new information to work from.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 7 objective 1 (diagnose underperforming prompts)Report a problem with this question

  15. 15. Claude describes an industry reporting requirement correctly as it stood some time ago, but the rule was amended recently. How should an associate classify and fix this?

    • A.As a context-window problem to be solved by starting a new conversation
    • B.As a formatting problem to be solved by asking for a cleaner layout
    • C.As a hallucination that proves the output cannot be trusted for any factual task
    • D.As a knowledge gap about recent change, not a fabrication: supply the current official text and have Claude work only from itAnswer

    A stale but genuine fact is a knowledge gap about recent developments, which is different from a hallucination (invented content) and from context loss (dropped earlier turns), and each has a different remedy. The remedy here is to supply the current authoritative text and restrict the model to it, which is also why anything asserted about very recent events warrants checking.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 2 (identify hallucinations vs other failure modes)Report a problem with this question

  16. 16. Late in a long working session, Claude's drafts stop honoring a plain-language rule the associate set at the start. What is the MOST appropriate response?

    • A.Conclude the model cannot follow style rules and stop using it for drafting
    • B.Accept the jargon-heavy drafts and edit each one by hand from now on
    • C.Treat it as constraint loss over a long conversation: restate the rule, or restart with the standing instructions and reference material carried forwardAnswer
    • D.Treat it as a hallucination and ask Claude to fact-check its own drafts

    An instruction that was honored early and quietly stops being applied points to constraint loss in a long conversation, not to fabricated content, so the fixes are restating, summarizing forward, or persisting the standing instructions. Fact-checking targets truth, not style adherence, and abandoning the workflow discards a task the model handles well once the constraint is re-established.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 3 (context limitations: when to restart, summarize or persist)Report a problem with this question

  17. 17. A communications associate has a Claude-drafted press release quoting a named customer and a specific savings percentage. It is scheduled to go out publicly tomorrow. What is the MOST appropriate step before release?

    • A.Soften the percentage to a range so it does not need to be checked
    • B.Ask Claude to confirm the quote and the percentage are accurate before publishing
    • C.Publish on schedule, since the release reads well and the quote is attributed
    • D.Route it for human review, confirming the quote with the named customer and the percentage against the underlying dataAnswer

    Public distribution, an attributed quotation, a named identifiable person and a hard number are all triggers for human review before release, because errors are irreversible once published. Asking the model to confirm its own output is not independent verification, and vaguing up a number hides rather than resolves an unverified claim.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 4 (determine when human review is required)Report a problem with this question

  18. 18. A department head asks Claude to decide which two of five team members should be moved off a struggling project. What is the MOST appropriate use of the output?

    • A.Implement the recommendation as given, since the model weighed the factors impartially
    • B.Ask Claude to rank all five people and act on the bottom two automatically
    • C.Use it only to structure the considerations, and keep the decision about individuals with the accountable managerAnswer
    • D.Ask Claude to justify its choice at greater length, then implement it

    Consequential judgments about identifiable individuals must be owned by an accountable person, so this is a poor fit for delegated decision-making even though it is a fine fit for organizing the considerations. A longer justification is not evidence of a sound decision, and automating the ranking removes the human oversight the situation requires.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 4 (when human review or ownership is required)Report a problem with this question

  19. 19. An analyst notices that a Claude report's opening summary says revenue grew 12 percent, while the table below it sums to roughly 8 percent. What is the BEST next action?

    • A.Keep the summary figure, since summaries are written after the analysis is complete
    • B.Treat the mismatch as an internal inconsistency and recompute both figures from the source data before using eitherAnswer
    • C.Ask Claude which of the two figures it prefers and use that one
    • D.Keep the table figure, since tables are inherently more reliable than prose

    A contradiction between the summary and the body is a classic inconsistency signal, and it tells you only that at least one figure is wrong, not which one. Resolution requires recomputing from the authoritative source data; neither section is inherently trustworthy, and asking the model to pick a favorite substitutes preference for evidence.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 2 (identify inconsistencies in responses)Report a problem with this question

  20. 20. A procurement associate asked Claude to recommend a software vendor, uploading only one vendor's own whitepaper as background. The recommendation strongly favors that vendor. What is the MOST appropriate concern?

    • A.The output is fine, since the reasoning is balanced in tone and cites the document
    • B.The output likely inherits the bias of its single source, so comparable material on the other vendors is needed before relying on itAnswer
    • C.The main risk is formatting, so the recommendation should be reissued as a comparison table
    • D.The output is fine, since it was grounded in a real document rather than invented

    Bias can enter through the supplied knowledge base itself: grounding in a single vendor-authored source produces a one-sided recommendation even when nothing is fabricated. A balanced tone is a presentation feature, not evidence of balanced inputs, and reformatting the same skewed content into a table does not correct the underlying imbalance.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 2 (identify biases, including inherited source bias)Report a problem with this question

  21. 21. An operations analyst needs Claude's comparison of 40 suppliers across six attributes, and the result will be loaded into a tracking spreadsheet and sorted. Which output format is MOST appropriate?

    • A.A narrative essay comparing the suppliers in prose paragraphs
    • B.A conversational inline answer, because it keeps the analyst in the flow of the chat
    • C.Whichever format produces the shortest response, to save reading time
    • D.Structured data such as a table or CSV, because the output feeds another system and must be sorted and filteredAnswer

    Format is chosen from what the consumer will do with the output, and here it will be ingested downstream and compared row by row, which is exactly what structured data supports. Prose and inline chat cannot be sorted or filtered reliably, and brevity is not a selection criterion.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 6 (select appropriate output formats)Report a problem with this question

  22. 22. A project manager will produce a substantial onboarding guide with Claude that her team must revise over several weeks and share with new hires. Which output choice is MOST appropriate, and why?

    • A.A series of short inline chat replies, because they are quicker to read as they arrive
    • B.An artifact, because the content is substantial and self-contained and needs to persist through versions, edits and sharingAnswer
    • C.Whatever format the first response happens to use, since format can be changed later anyway
    • D.Structured data, because any long document is easier to manage as rows and columns

    Artifacts are meant for substantial, self-contained content that the user will keep working on, since they persist alongside the conversation and support versioning, editing and sharing. Inline replies suit short conversational answers, structured data suits row-wise comparison or downstream ingestion, and drifting into whatever format appears first is not a decision at all.

    Source: Anthropic Claude certification (Associate – Foundations), Domain 2 objective 6 (artifacts vs inline vs structured data)Report a problem with this question

Practice questions based on the official Claude Certified Associate – Foundations (CCAO-F) exam guide and Anthropic's public documentation. This is an independent study tool, not affiliated with or endorsed by Anthropic, and does not grant certification. The real exam is 60 questions, 120 minutes, passing at a scaled 720/1000, delivered via Pearson VUE ($99). Official certification page →