Fixed-corpus descriptive study
Criterion-level quality ratings across four listening operations
Thirty-two first-generation items are planned: eight each for Gist, Detail, Inference, and Speaker Attitude/Intention. Two raters review every item independently using eight YES/NO/UNCLEAR criteria.
Loading current results…
RQ1
Criterion-level distributions
RQ2
Patterns by listening operation
Context check
Source × operation raw counts
Cells contain four judgments per criterion when complete. Counts are shown instead of tiny-cell percentages.
Method QA
Independent-rater agreement
Raw agreement and unweighted Gwet’s AC1 are primary reporting outputs. Unweighted Cohen’s κ is shown only as a secondary comparison.
Interpretation boundaries
What these results can and cannot show
- Percentages describe this fixed 32-item corpus; they are not population prevalence estimates.
- The study does not establish formal CEFR validation, empirical difficulty, discrimination, or empirical distractor functioning.
- Operation and recording comparisons are descriptive. They do not estimate source effects or causal effects.
- Agreement describes consistency in applying the checklist. It is not test reliability or proof that the checklist is valid.