What Is Pre-Employment Testing? A Practical Guide to Choosing and Using Assessments

Pre-employment assessments are tests used to evaluate abilities, personality, knowledge, and behavior that may be relevant to a candidate’s future work.

An assessment is not inherently “good” or “bad” on its own. Whether it is appropriate depends on how it is connected to the actual work in a particular company and role.
Starting with “choose a famous test” is therefore almost automatically the wrong approach.

Scores are only one part of a broader selection model

Common selection methods can be grouped roughly as follows. In everyday usage, the non-interview methods are typically treated as pre-employment assessments.

Method Main information obtained Best use
Interview Judgment, experience, and behavior in work-related situations Broad coverage of the role
Personality assessment (Big Five) Relatively stable behavioral tendencies Building and testing hypotheses about work behavior
Personality assessment (proprietary model) Response patterns expressed through proprietary labels Generating hypotheses to check in interviews
General cognitive ability test Abstract reading, numerical, and reasoning tasks Checking minimum required levels
Work sample Performance on tasks similar to the actual job Direct assessment of major tasks

These methods should be combined into a selection process suited to each company and role.
The most important objective is to evaluate the actual work from multiple angles rather than treating any single score as a complete answer. Even when an assessment is useful, it is still only one component of the overall selection process.

A practical starting point is to make sure the interview covers the important parts of the role that the available assessments do not.
For example, if handling customer complaints is important, there may be no off-the-shelf assessment that directly captures the full situation. The interview should therefore include questions that cover complaint handling rather than leaving that area unexamined.

Compare pre-employment assessment services

Keep the interview, but structure it

The first rule for interviews is to stop treating them as free-form conversations.
When the topics differ from candidate to candidate, judgments are easily pulled by first impressions, résumé details, or whatever happens to come up in the conversation.

For important parts of the job, it is useful to:

  • prepare common questions for all candidates
  • define evaluation criteria in advance
  • use interview questions to cover areas that assessments do not capture

A 2022 meta-analytic reexamination by Sackett and colleagues found that many long-standing validity estimates for selection methods had been overstated, while structured interviews remained among the strongest methods in the comparison.1

Personality assessments are not pass/fail tests

Unlike many ability tests, personality assessments do not usually have a single direction in which “higher is always better.”

For example, higher extraversion may be helpful in some customer-facing situations, but the same direction may be less important in work that requires prolonged individual concentration. Traits such as conscientiousness are related to performance across many jobs, but their practical meaning still depends on the work being performed.

The Big Five has an important advantage: it provides a relatively common vocabulary for personality across different tests and research traditions. Proprietary personality systems require more caution, because the labels shown to users may not map cleanly onto well-established psychological constructs. Their practical meaning therefore has to be checked against observed work behavior.

Self-report personality scores can also be influenced by self-presentation and response style. It is safer to treat the result as a hypothesis to compare with interviews and other evidence rather than as a hidden “truth” about the person.

General cognitive ability tests no longer support “higher is always better”

General cognitive ability (GMA) tests have traditionally been treated as strong predictors under the assumption that higher scores imply proportionally better future job performance.

That interpretation has been substantially weakened by newer evidence.

The idea that a candidate with a score equivalent to the 60th percentile should therefore be expected to outperform one at the 55th percentile is no longer a defensible general rule.

Berry, Lievens, Zhang, and Sackett (2024) used an updated meta-analytic matrix of multiple selection methods and found that removing GMA tests from a selection system generally produced little or no loss in overall validity while substantially reducing racially adverse selection outcomes.2

This work followed Sackett et al. (2022), who reexamined range-restriction corrections in personnel-selection research and reduced the estimated correlation between GMA and job performance from the classic estimate of .51 to .31.

There is still a reasonable case for using ability tests to identify clear low-end outliers who may not meet a genuine minimum requirement of the job. Multi-stage selection with minimum thresholds is not inherently problematic, but the evidence does not support treating broad central score ranges as a precise ranking of future employees.

If this makes the practical value of many general ability tests seem limited, that is a reasonable reading of the current evidence.

Examples of a more defensible use include:

  • checking basic reading comprehension in work that requires processing large volumes of written information accurately
  • checking basic numerical competence in work where frequent numerical errors would create material problems

Because broad score ranges have limited explanatory value, it is also difficult to argue that simply adding a GMA score to an interview necessarily produces a meaningfully better prediction model.

Work samples make the connection to the job easier to inspect

A work sample asks candidates to perform tasks that resemble the actual work, making the connection between the test content and the job easier to inspect directly.

Examples include:

  • asking an engineer to write a short piece of code
  • using a live conversation task for a role that requires customer communication in another language
  • simulating customer discovery and proposal work for a sales role
  • asking an administrative candidate to perform a realistic checking or reconciliation task

However, “looks like work” is not enough. Most jobs contain many different tasks. If a selection exercise represents only one narrow slice of the job, it becomes harder to justify conclusions about overall future performance.

The set of work-sample tasks should therefore represent the role broadly enough to avoid over-weighting one visible skill. This requires at least some analysis of what the job actually contains.

For example, spoken English may matter in a role without constituting most of the job. A realistic approach is to estimate how much weight that activity deserves and evaluate the remaining work through interviews or other methods.

The main practical limitation is availability. Off-the-shelf work-sample services exist for areas such as software engineering, but many roles require an organization to create its own simulations or role plays.

One important caution is that multiple-choice tests of job knowledge should not be treated as work samples. Real jobs are rarely performed by selecting answers from a quiz interface.
Because such tests can have low directness and narrow coverage, they are better treated as weak supporting information for interviews rather than as direct representations of job performance.

Do not let one assessment decide the entire hiring outcome

The quality of a selection process depends not only on individual tests but also on how multiple methods are combined.

Two broad approaches are common.

Multiple hurdles: check minimum requirements in sequence

Each stage has a minimum requirement. Candidates who do not meet it do not proceed.

For example:

  1. mandatory qualification or experience
  2. structured interview
  3. work sample

This is easy to understand when a requirement is genuinely necessary, but unnecessary cutoffs can eliminate otherwise strong candidates.

Compensatory evaluation: combine different strengths

Multiple pieces of evidence are combined, allowing one strength to compensate for another weakness.

For a customer-facing role, for example, the evaluation might combine:

  • communication accuracy
  • problem solving
  • conscientiousness
  • judgment shown in structured interview questions

In practice, a hybrid is often useful: use hurdles only for truly mandatory requirements, then use a broader combined evaluation afterward.

The important point is not to adopt a vendor’s overall score as the hiring rule. Each company still has to decide what should be a hard requirement and what should be interpreted together with other evidence.

Validity is evidence for the inference you make, not the popularity of the test

A selection method is useful only when there is a clear reason to believe that its score supports a work-related inference. Two forms of evidence are especially important:

  • Relationships with other variables: whether test scores relate to job performance or other relevant outcomes in the expected way
  • Relationship to job content: whether the assessment tasks adequately represent the work being evaluated

Vendors often advertise numbers such as “companies using the test,” “candidates assessed,” or “millions of data points.” Those figures are not, by themselves, evidence that the assessment supports a particular hiring decision.

Internal structure—whether items actually form the psychological dimensions the test claims to measure—and response processes—whether people answer through the intended cognitive or judgment process—are useful supporting evidence, but they are weaker on their own because they do not establish a connection to the job.

Refine your understanding of the role as you use the selection process

In practice, organizations rarely begin with a complete and reliable model of every requirement in a role.

A more realistic starting point is to identify what the current interview is already trying to check, what has caused problems after hiring, and which important areas are being missed.

For example, if complaint handling matters, “interpersonal ability” is too abstract to be useful by itself. It can be broken into more observable activities such as fact-finding, responding to an upset customer, proposing a resolution, and escalating the case appropriately.

That makes it easier to see what should be checked in an interview, what might be tested through a work sample, and where a personality assessment may provide useful hypotheses.

Assessments can therefore help refine hiring criteria, but they do not automatically become the criteria themselves.
The practical goal is to begin with a rough understanding of the work, use additional evidence to improve that understanding, and revise the selection process as better information becomes available.

A mechanical score does not automatically make hiring fair

Fair hiring requires organizations to base decisions on characteristics that are meaningfully related to the work rather than on race, ethnicity, social background, disability, or other irrelevant characteristics.

Standardized assessments can reduce some sources of inconsistency because all candidates receive the same procedure.

But mechanical scoring does not guarantee fairness. The entire combination of interviews, assessments, and decision rules still has to relate sensibly to future performance in the actual work.

At minimum, fairness requires attention to questions such as:

  • whether candidates receive comparable testing conditions
  • whether access to instructions and practice opportunities differs
  • whether predictions are systematically too high or too low for particular groups

As discussed above, cognitive ability tests are known to produce substantial average score differences across demographic groups. In many selection systems, removing them can improve fairness at little or no loss in predictive validity.

When enough hiring data are available, organizations should examine not only relationships with post-hire performance but also major differences in pass rates and prediction errors across candidate groups.

A practical way to introduce pre-employment assessments

A realistic implementation sequence is:

  1. Start from the current interview
    List what interviewers are already trying to assess, where hiring mistakes have occurred, and what important areas are being missed.

  2. Make sure the interview covers the role broadly
    Keep important areas in structured interview questions when no suitable assessment exists.

  3. Prefer direct assessment where practical
    Use work samples or other methods that closely resemble the actual job when they are available.

  4. Use Big Five personality assessments to add a stable descriptive framework
    Use well-established trait dimensions to generate work-related hypotheses rather than treating personality scores as automatic hiring criteria.

  5. Decide how methods will be combined
    Separate true pass/fail requirements from information that should contribute to an overall judgment.

  6. Review results and update the model
    Compare selection information with post-hire outcomes and refine questions, criteria, and assessments over time.

The organization and the job will also change. A selection model that was once reasonable may become less useful as responsibilities, technology, or business conditions shift. The goal is not to create criteria that never change, but to avoid treating old evidence as permanently valid when the work itself has changed.

What to check when comparing assessment services

When choosing a service, focus less on brand recognition or the number of questions and more on:

  • what the assessment claims to measure
  • whether research or technical documentation supports that claim
  • how the measurement connects to the actual work
  • whether the score is intended for ranking, a minimum threshold, or interview support
  • how it can be combined with other selection methods
  • whether candidate instructions and testing conditions are clearly managed
  • whether the method can be revised when the job or hiring model changes

Compare pre-employment assessment services

References

⁋ Jun 30, 2026↻ Sep 8, 2026