Using GPT-6 Astra to Screen Candidates: The Workflow, the Prompts, and the Guardrails
Otavo's recruiting team runs applicants through GPT-6 Astra. This is the workflow we use: fixed scorecards, quoted evidence, and prompt patterns you can copy, plus the monitorability and audit-trail problem it creates and the compliance floor we hold to.

Otavo's screening workflow ranks candidates against a fixed scorecard, with a human reviewing every call.
How is GPT-6 Astra used to screen job candidates?
Write a weighted scorecard before you screen, then run every applicant through one unchanged prompt that returns structured, per-criterion output. Require a quoted line of evidence for each rating, have Astra flag gaps rather than guess, and keep every rejection with a human reviewer.
Key takeaways
- Write the weighted scorecard before you touch the model; the rubric decides the outcome more than the model does.
- Run every applicant through one unchanged prompt that returns structured, per-criterion output, and require a quoted line of evidence for each rating.
- Tell the model that flagging missing information is a correct answer, and keep every rejection decision with a human reviewer.
- Unilever's early-careers program reported roughly a 75% cut in time-to-hire, but from pre-frontier assessment tooling rather than anything like Astra.
- SHRM data shows time-to-hire has risen even as AI adoption reached roughly 43% of HR teams in 2026, and Gartner reports only about 26% of applicants trust AI to judge them fairly.
The applicant pile has stopped being readable. Harvard Business Review argued in June 2026 that AI has broken the hiring funnel at both ends: candidates now use the same frontier models to fire off hundreds of tailored applications, and recruiters answer with AI screening, which is why Robert Half reports that 67% of HR leaders see time-to-hire going up rather than down despite the tooling. That is the paradox worth sitting with. A resume polished by a model is no longer a reliable first-round signal, and the volume arriving on Monday morning no longer fits inside a human week.
What follows is the screening workflow Otavo actually runs: the scorecard we write before anything goes near a model, the prompt patterns we paste in verbatim, and the ranking pass that keeps the two from drifting apart. It also sets out what we refuse to let the model decide, which is a shorter list than the workflow and a more important one.
Screening is one of the few places where a fluent summary can quietly become a decision about someone's livelihood, so the harder parts get their own sections: where Astra's reasoning has become difficult to inspect, and what the law already asks of employers who automate selection. One naming note before the workflow: Astra, which shipped on September 3, 2026, is simply the name of the GPT-6 release, one model, the way Sol names GPT-5.6.
What Astra actually changes for screening work
Everything in this section is vendor-reported, drawn from OpenAI's evaluations of its own model. The figures are still useful, because they describe capabilities that map onto screening work.
Long-context retrieval that holds its shape
Screening is a retrieval problem before it is a judgment problem. A recruiter reading a six-page resume and a cover letter is hunting for a few signals buried in formatting noise. OpenAI reports Astra scoring 100% on MRCR v2 8-needle retrieval at 256K to 512K tokens and 96.3% at 512K to 1M tokens, against 91.5% and 73.8% for GPT-5.6 Sol at the same context lengths. That is the difference between a model that reads a full batch evenly and one that gets vague about the first few profiles.
A lower hallucination rate
On OpenAI's internal hallucination benchmark, where lower is better, the company reports 4.2% for Astra against 12.2% for GPT-5.6 Sol. OpenAI also reports that Astra is roughly three times less likely than Sol to make inaccurate claims about its own capabilities. That second figure may matter more for hiring: a model that overstates what it can do will confidently rate a criterion it had no evidence for.
Trained for professional deliverables
OpenAI describes Astra as trained for professional work: following an existing template rather than reinventing the format, producing structured documents, spreadsheets, and presentations, and pulling only the relevant context into an output. A screening pass lives or dies on this. Return the same fields, in the same order, with the same allowed values across 40 profiles, and the pass is reviewable weeks later. Return essays and it is not.
Asking rather than assuming
One design behavior is worth building your process around. OpenAI says Astra fills routine gaps from context on its own while asking a focused question when the answer could materially change the outcome. That is close to what a good screener does with a thin resume. Tell the model that flagging a gap is a correct answer, and it will take the offer.
A screening workflow that holds up
Most disappointing results come from pasting in a resume and asking whether the person is a good fit. That question invites the model to invent a standard and then judge against it. The sequence below removes the invention.
- Write the scorecard before you touch the model. Agree with the hiring manager on what you are assessing before any profile goes near Astra. If it is not written down, the model will supply one, and you will screen against its assumptions.
- Convert the job description into weighted criteria. Four to six criteria, each with a weight and a short description of what strong evidence looks like. Weights force the conversation nobody wants: what you would trade away to get what matters most.
- Strip identifying details before the model sees a profile. Remove names, photos, addresses, and personal contact details. It is partly privacy and partly quality, because it denies the model signals you would not want a human screener leaning on either.
- Batch profiles for consistency. Screen in batches of roughly five to eight against one unchanged prompt. A consistent standard across the batch beats depth on any single candidate, and drifting standards are the fastest route to an unfair pass.
- Demand evidence with a quote, not a verdict. For every criterion, require a rating plus the line from the profile that supports it. Quotes are checkable in seconds. A verdict you cannot check is one you have quietly delegated.
- Require the model to flag missing information rather than guess. Give it a category such as no evidence found, and state that using it is the right answer when the profile is silent. Half of good screening is knowing what to ask on the phone.
What comes out is not a shortlist. It is a structured reading of each profile against agreed criteria with the gaps marked, which beats a confident paragraph about culture fit.
Ranking profiles defensibly
Ranking is where teams get seduced. A single number per candidate looks tidy in a spreadsheet, but it compresses unrelated judgments into one figure nobody can take apart afterward. Four practices keep ranking useful and one keeps it defensible.
- Score per criterion, never overall. Keep the rubric ratings visible and apply your weights yourself, in a sheet you control. Then the answer to why one candidate sits above another is a named criterion, not a vibe.
- Use forced-choice pairwise comparison for the shortlist. Asking which of two candidates better meets the top criterion, and why, produces sharper reasoning than a ranked list of eight. Pairwise only scales to a shortlist, which is where precision matters anyway.
- Calibrate against known-good past hires. Run anonymized profiles of people who went on to succeed in the role through your prompt. If your strongest performers come out middling, your criteria are wrong, not your candidates.
- Run the same batch twice. Same profiles, same prompt, fresh conversation, then compare. Ratings that move between runs tell you the criterion is too vague to score, and you want that found before it reaches a candidate.
- Preserve the outputs as a record. Save the prompt version, the criteria, the model version, the structured output, and the quotes alongside the requisition. If anyone asks how your process treated a group of applicants, the useful answer is a paper trail.
The monitorability problem
Hiring compliance runs on one practical ability: explaining, later, why a particular candidate was screened out. Not a score, but the criterion, the evidence, and the person who made the call. Everything else in an AI-assisted process serves that requirement.
Which is why the most significant thing OpenAI published about Astra is not a benchmark. In its own evaluations, OpenAI found Astra's written reasoning harder to monitor than GPT-5.6 Sol's. It attributes the decline to Astra solving problems in fewer written steps, says it takes the finding seriously, and calls monitorability an ongoing research priority. Independent reporting raises a related concern, that a new reasoning approach obscures part of the chain of thought.
Be fair about the disclosure. OpenAI surfaced this itself when nothing compelled it to, which is what you want from a vendor sitting in the middle of consequential decisions. It does not make the problem go away, though. It hands it to you.
Two other claims sit awkwardly together. OpenAI describes Astra as its most aligned model to date, while also stating that it meets the Critical threshold for cybersecurity capability under OpenAI's Preparedness Framework. Neither is about hiring, and neither tells you whether a screening output can be explained to a regulator.
If your defense of a screening decision depends on the model's reasoning trace, you do not have a defense. You have a transcript the vendor has already told you is harder to monitor than the last one.
So build the audit trail out of artifacts you control: the lines the model quoted from each profile against each criterion, the criteria, the prompt version, the model version, and your reviewer's decision with a short written reason. Those are stable, human-readable, and defensible. The model's internal reasoning is none of the three.
Prompt patterns you can adapt tomorrow
Replace the bracketed parts, keep the rules, and resist the urge to trim the constraints. The constraints are what make the output reviewable.
ROLE: [job title]
MODEL: GPT-6 Astra (record the model version with the output)
You are helping with a first-pass screen.
Do not recommend, reject, or rank anyone.
CRITERIA AND WEIGHTS:
1. [criterion] (weight 30) strong evidence looks like: [...]
2. [criterion] (weight 25) strong evidence looks like: [...]
3. [criterion] (weight 25) strong evidence looks like: [...]
4. [criterion] (weight 20) strong evidence looks like: [...]
PROFILE (identifying details removed):
[paste anonymized profile]
For each criterion, return:
- rating: strong | adequate | weak | no evidence
- evidence: a direct quote from the profile, or "none found"
- confidence: high | medium | low
Then return:
- missing_information: what you would need in order to rate this fairly
- questions_for_screen_call: up to 3, tied to the weakest evidence
Rules:
- Use only what the profile states. Do not infer a rating.
- If the profile is silent on a criterion, answer "no evidence".
- Ask me a clarifying question only if the answer would materially
change a rating. Otherwise proceed and flag the gap.
- Ignore names, schools, employer prestige, photos, employment gaps,
and anything else unrelated to the criteria.
- Return JSON with exactly the keys above, in the same order for
every profile in this batch.Once a batch has been screened, the ranking pass reuses the same criteria so the two stages cannot drift apart.
You are comparing anonymized candidates for one role.
Output data and evidence only. No hiring recommendation.
CRITERIA AND WEIGHTS: [paste the same list used for screening]
CANDIDATES: [paste 5 to 8 anonymized profiles, labeled A, B, C, ...]
Step 1: For every candidate, rate each criterion
strong | adequate | weak | no evidence,
each with a supporting quote.
Step 2: Return a table of candidates by criterion.
Do not compute an overall score.
Step 3: For each pair of candidates, say which one better meets the
highest weighted criterion and why, in one sentence that
cites evidence.
Step 4: List any candidate whose ratings would change materially if
one piece of missing information turned out differently.
Rules:
- Use only what the profiles state.
- Flag criteria you could not assess instead of filling the gap.
- Do not produce a final ranking, and do not recommend advancing
or rejecting anyone.
- Keep the output in the same structure across reruns of this batch.Keep both prompts in a shared document with a version number. The moment two recruiters run different prompts on the same requisition, you no longer have a process, you have two.
What Astra should not decide
The line worth defending runs between structuring information and deciding outcomes. It is easy to cross by accident when the output is fluent and the pile is large.
- A person reviews before any rejection. No automatic filtering, no bulk rejection driven by a score. If the volume makes real review impossible, the honest fix is a narrower funnel, not a faster filter.
- Tell candidates, and mean it. Say plainly where AI assists your process and where humans decide. Handle notice, consent, and any request for an alternative route with your legal and privacy colleagues rather than improvising per candidate.
- Candidate data never goes into a consumer account. Use approved enterprise tooling with your organization's data terms, retention settings, and access controls. Zero Data Retention is supported for eligible API customers and is worth pursuing before the first batch, not after.
- Treat access as a decision, not a default. Enterprise administrators must enable Astra for their workspace, and access is off by default at launch. Use that gap to agree the guardrails before recruiters experiment on live applicants.
The access facts as of launch: Astra is available to ChatGPT Plus, Pro, Business, and Enterprise users, through the OpenAI API as gpt-6-astra, and through Microsoft Azure and AWS Bedrock. API standard pricing is $10 per million input tokens and $50 per million output tokens, with a Fast mode offering up to 2x the speed at 2x the price. For batch screening, standard mode is almost always right.
How to start this week
- Confirm Astra is enabled for your workspace under terms that cover candidate data, and set your retention posture before anything is pasted in.
- Pick one open role and write a four to six criterion scorecard with weights, agreed with the hiring manager.
- Calibrate by running anonymized profiles of past strong performers through the screening prompt, then fix the criteria that misfire.
- Screen one batch of five to eight anonymized profiles, twice, and compare the runs for stability.
- Review every output before anyone is advanced or rejected, and record where you disagreed with the model and why.
- Ask counsel which regimes apply to your locations, and add a line to your careers page describing how AI assists your process.
None of this requires a platform migration. It requires deciding what you are measuring, keeping the model on evidence rather than opinion, and writing down the reasoning the model can no longer be relied on to show you. The teams getting real value from Astra are not the ones asking it who to hire. They are the ones who arrive at the debrief having read all three hundred applications.
Frequently asked questions
Sources
- GPT-6 Astra: A new generation of intelligence — OpenAI (September 3, 2026)
- GPT-6 Astra System Card — OpenAI (September 3, 2026)
- Safety overview: GPT-6 Astra — OpenAI (September 3, 2026)
- OpenAI launches new Astra model amid growing scrutiny over agents' safety — Reuters (September 3, 2026)
- OpenAI's new reasoning technique alarms AI safety experts — TechCrunch (September 2, 2026)
- Automated Employment Decision Tools (AEDT), Local Law 144 of 2021 — NYC Department of Consumer and Worker Protection
- Annex III: High-Risk AI Systems Referred to in Article 6(2) — EU Artificial Intelligence Act
- Select Issues: Assessing Adverse Impact in Software, Algorithms, and Artificial Intelligence Used in Employment Selection Procedures Under Title VII of the Civil Rights Act of 1964 — U.S. Equal Employment Opportunity Commission (May 18, 2023)
- The Americans with Disabilities Act and the Use of Software, Algorithms, and Artificial Intelligence to Assess Job Applicants and Employees — U.S. Equal Employment Opportunity Commission (May 12, 2022)
- Amazon scraps secret AI recruiting tool that showed bias against women — Reuters (October 10, 2018)
- AI Has Broken Hiring. Here's How to Fix It. — Harvard Business Review (June 1, 2026)
- AI in Recruiting: Why Hiring is Harder in 2026 — Robert Half
- I Tried OpenAI's Astra-Powered GPT-6 Pro For Job-Search — Forbes (September 6, 2026)
Fact-checked by Nilesh Kapoor on September 11, 2026