Skill Validation

Communication Skills: Selection Validity Evidence and Measurement Methods

By Editorial Team — reviewed for accuracy Published
Last reviewed:

Communication skills are the most-cited and least-rigorously-measured construct in selection practice. Almost every job description mentions “strong communication”; almost no hiring loops measure it with the rigor applied to coding tests or cognitive ability assessments. The asymmetry is partly historical (verbal selection tests are harder to standardize than written ones) and partly construct-level (communication is genuinely multi-dimensional, and most single tests probe only one dimension). The peer-reviewed literature, however, supports communication as a measurable construct with validity-defensible assessment formats. Hunter, Schmidt, and Judiesch (1990) reported that communication-related work samples and structured interviews carry meaningful incremental validity over cognitive-ability tests, and Schmidt and Hunter’s 1998 meta-analytic synthesis sets the work-sample ceiling at ~0.54 corrected operational validity.

This article walks through the construct decomposition, the validity evidence base, AIEH’s communication sample and how it integrates into the Skills Passport composite, the recurring measurement pitfalls, and the takeaway for teams designing communication assessment into the loop.

Data Notice: Validity coefficients cited here reflect peer-reviewed meta-analytic evidence at time of writing. Specific weights AIEH applies to communication evidence are documented in the scoring methodology and may evolve as calibration data accrues.

What communication assessment measures

The communication construct decomposes into multiple layers, and a defensible assessment is explicit about which it samples:

  • Written clarity. Given a complex idea or set of facts, can the candidate produce written prose that is well-organized, accurate, and tonally appropriate for the audience?
  • Verbal articulation. Given a question or topic, can the candidate explain their reasoning verbally with logical structure, appropriate pacing, and responsive to audience cues?
  • Audience adaptation. Can the candidate adjust register, vocabulary, and depth based on the audience — technical for engineers, business-frame for executives, plain-language for end users?
  • Listening and synthesis. Can the candidate receive complex input, identify the underlying question, and respond to that question rather than to a misread of it?
  • Conflict and difficult-conversation handling. Can the candidate navigate disagreement, deliver hard feedback, and de-escalate without losing signal?

Argyle’s Bodily Communication established that the verbal channel carries only a fraction of the total communicative payload — a finding that complicates single-channel assessment but doesn’t undermine the construct. The implication for selection is that a defensible communication assessment should sample multiple channels and dimensions rather than treating “communication” as a single number.

The validity evidence base

Hunter, Schmidt, and Judiesch (1990) reported incremental validity for communication-related predictors above cognitive-ability tests, with the effect sizes varying by job complexity — communication matters more for higher-complexity roles where the job involves substantial coordination. Schmidt and Hunter’s 1998 meta-analytic synthesis set the work-sample ceiling at ~0.54 corrected operational validity, and a communication work sample (a written draft, a recorded explanation, a customer-email response task) inherits that ceiling when the design preserves point-to-point correspondence with the criterion.

The structured-interview literature is also relevant. Schmidt and Hunter treated structured interviews as delivering corrected operational validity around ~0.51, and a structured interview that scores communication-relevant rubric dimensions (organization, clarity, audience adaptation) inherits that ceiling. McDaniel et al (1994)‘s meta-analytic review of employment interviews reached similar conclusions about the validity gain from structure.

Sackett and Lievens (2008) frame the predictor-criterion fidelity question as the core validity driver. A communication assessment that asks the candidate to perform a representative slice of the actual communicative work the role does inherits the validity ceiling; an assessment that asks “rate yourself on communication” sits well below.

The empirical literature on AI-mediated communication assessment is younger but growing — particularly relevant as AI tools change how candidates draft, revise, and deliver communicative output.

AIEH communication test integration

AIEH treats communication evidence as a primary input to the communication pillar of the Skills Passport composite. The communication sample presents candidates with realistic communicative tasks — drafting an email to a frustrated customer, explaining a technical concept to a non-technical audience, summarizing a meeting transcript — and scores the output against an anchored rubric covering clarity, organization, accuracy, tone, and audience fit.

The composite weights communication evidence higher for roles where coordination is diagnostic — customer-facing, managerial, cross-functional — and lower for heads-down individual-contributor roles. The scoring methodology documents default weights and role-bundle modifiers.

For hiring teams, the hire workspace surfaces communication evidence as a sub-score within the communication pillar, with provenance visible alongside the calibrated number. The skills-based hiring evidence article covers the broader frame for why portable skill evidence outperforms credentials.

AI-mediated communication: a new measurement question

A recent complication in communication assessment is the rise of AI-mediated drafting. A candidate producing written prose with LLM assistance is demonstrating a different construct than one producing prose unaided — partly the underlying communicative skill, partly the AI-collaboration skill of using a model to generate, evaluate, and refine output. For roles where the on-job criterion involves AI-mediated communication (most contemporary knowledge work), the AI-aided assessment is the more diagnostic predictor.

This shifts the design question. A communication assessment administered in a clean-room environment without AI tools measures a construct partially diverging from the criterion for any role where the candidate will use such tools daily. The AIEH ACL prompt-to-spec sample captures the AI-collaboration side specifically, and hiring teams should consider whether their communication assessment should include both the unaided and the AI-aided variants.

The construct decomposition holds in either mode — clarity, organization, accuracy, tone, audience-fit remain the rubric dimensions. What changes is the process the candidate uses to produce the output, and the rubric should be agnostic to process while being specific about output quality. The ai fluency in hiring article covers the broader frame for combining traditional and AI-aware skill measurement.

Pitfalls that collapse communication test validity

Several recurring failure modes drop a communication assessment’s predictive power below the work-sample ceiling:

  • Self-report items. “Rate your communication skills” produces almost no usable signal. Self-report scales correlate weakly with observed communicative performance.
  • Mono-channel assessment. A test that samples only written prose predicts a different criterion than one sampling verbal articulation. For roles where both matter, both should be sampled.
  • Unanchored rubrics. “Did the candidate communicate well?” produces wide rater variance. A defensible rubric decomposes into clarity, organization, accuracy, tone, and audience-fit dimensions with anchored examples at each band.
  • Single rater. Two raters scoring independently and reconciling produces substantially lower variance than single-rater scoring, particularly for the more subjective dimensions like tone.
  • Pattern-matching on accent or register. Interviewers scoring “communication” sometimes pattern-match on accent, dialect, or sociolinguistic register rather than on the construct dimensions. A defensible rubric scores artifacts and clarity, not delivery style.
  • Mismatched audience. A communication task scored by raters who aren’t the target audience measures a different construct than one where raters represent the actual audience. Customer-email tasks should be scored against customer-perception rubrics.

For practical guidance on integrating communication evidence with broader interview signal, see the structured interview design and interview question design articles.

Takeaway

Communication is a measurable construct with work-sample-grade validity (~0.54 corrected operational validity) when assessed via realistic communicative tasks scored against anchored rubrics by multiple raters. Validity collapses when assessments rely on self-report, sample only one channel, use unanchored rubrics, or pattern-match on accent and delivery style. AIEH’s communication sample feeds into the communication pillar of the Skills Passport composite at high relevance for coordination-heavy roles.

For hiring teams, the practical implication is that the “strong communication” line in every job description deserves an assessment with the same rigor applied to technical skill testing. A communication score generated from an anchored work sample is a substantially better signal than the impressionistic “this person communicates well” comment that often serves as the only communication evidence in unstructured loops. The hiring loop design article covers the architecture for combining communication evidence with cognitive, AI fluency, and domain signal, and the hiring bias mitigation article covers the rubric design that reduces accent- and register-bias contamination of communication scores.

Sources

  • Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274.
  • Sackett, P. R., & Lievens, F. (2008). Personnel selection. Annual Review of Psychology, 59, 419-450.
  • Hunter, J. E., Schmidt, F. L., & Judiesch, M. K. (1990). Individual differences in output variability as a function of job complexity. Journal of Applied Psychology, 75(1), 28-42.
  • Argyle, M. (1988). Bodily Communication (2nd ed.). Methuen.
  • McDaniel, M. A., Whetzel, D. L., Schmidt, F. L., & Maurer, S. D. (1994). The validity of employment interviews: A comprehensive review and meta-analysis. Journal of Applied Psychology, 79(4), 599-616.
  • Schmidt, F. L., Oh, I.-S., & Shaffer, J. A. (2016). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 100 years of research findings. Working paper.

About This Article

Researched and written by the AIEH editorial team using official sources. This article is for informational purposes only and does not constitute professional advice.

Last reviewed: · Editorial policy · Report an error