Public, de-identified reporting

AI-generated C1-targeted listening MCQs

Fixed-corpus descriptive study

Criterion-level quality ratings across four listening operations

Thirty-two first-generation items are planned: eight each for Gist, Detail, Inference, and Speaker Attitude/Intention. Two raters review every item independently using eight YES/NO/UNCLEAR criteria.

Loading current results…

RQ1

Criterion-level distributions

RQ2

Patterns by listening operation

Context check

Source × operation raw counts

Cells contain four judgments per criterion when complete. Counts are shown instead of tiny-cell percentages.

Method QA

Independent-rater agreement

Raw agreement and unweighted Gwet’s AC1 are primary reporting outputs. Unweighted Cohen’s κ is shown only as a secondary comparison.

Interpretation boundaries

What these results can and cannot show