Ask what the feature does when it is uncertain, how it was evaluated, and what a recruiter sees behind a score. Genuine capability comes with specifics: inputs used, known limitations, measured behaviour. Marketing language describes outcomes without mechanism. The fastest test is running the feature on your own data and comparing it with your own judgement.
Ask what happens when the system is uncertain. A real implementation has a defined behaviour: a lower confidence score, a flag for review, a fallback to keyword matching. Marketing has no answer because the question was never considered. Ask what the feature does badly, and treat a claim of no limitations as a finding. Ask what inputs it uses and what it deliberately ignores. Ask how it was evaluated, against what data, and what the result was. Ask what a recruiter sees when they open a ranked candidate, and whether the explanation reflects actual inputs. Vendors who built the capability answer these fluently and often volunteer caveats. Vendors who assembled a feature quickly tend to redirect to benefits, which is the tell.
Run the feature on your data rather than watching it on theirs. Bring resumes from a role you filled and see whether the ranking resembles your own conclusions. Include awkward cases: a career changer, an unconventional format, a strong candidate with a gap, a plausible but unsuitable applicant. Watch for keyword behaviour dressed as comprehension, which shows up when a candidate ranks highly purely for repeating the job title. Ask the presenter to change something live, such as adjusting the criteria, and see whether the output responds sensibly. Also notice whether the feature is presented as assisting a recruiter or replacing one, since the second framing usually accompanies weaker evidence. This is the same hands-on approach that separates [genuinely useful AI-assisted tools](/ai-recruiting-tools) from repackaged filters.
Predictions about future job performance from application data, which is difficult to validate and rarely accompanied by evidence you can inspect. Percentage improvements quoted without a described methodology or a named source. Claims of eliminating bias, which overstates what any system can guarantee and is a stronger statement than most vendors can support. Descriptions of a feature as fully automatic when a human is in fact reviewing output, or the reverse, where automation is understated. And roadmap capability presented in the present tense. None of these means a product is poor. They mean the marketing has run ahead of the evidence, and your job during evaluation is to establish which claims you can verify yourself and which you are being asked to accept.
Less than it sounds, provided it works. A well-built rules engine that ranks applicants usefully is a perfectly good product, and the label is irrelevant to your outcomes. The reason to distinguish them is expectation and price. If you are paying a premium for a tier on the basis of intelligence that turns out to be keyword filtering, you have overpaid for something the base product may already do. Judge the feature on measured behaviour against your data, then judge the price against that behaviour. Keep your assessment tied to whether recruiter time falls and shortlist quality holds, which is the only question your hiring outcomes care about.
Get a personalized walkthrough of Pitch N Hire on your own roles and workflow. No slides, no obligation.
Prefer to talk? Book a demo · View pricing
Free 1-user plan · No credit card · Talk to a real hiring expert
See your true cost-per-hire and how much Pitch N Hire could save you — our free Recruitment ROI Calculator gives you the numbers in under a minute. No signup required.
Open the free ROI calculatorPrefer a tailored walkthrough on your real roles? Drop your work email:
★ Free 1-user plan · No spam · Talk to a real hiring expert