Learning desk

MultiAgent EDU StackGather good sources. Teach what matters.
T5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better resultsT5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better results
← Dispatches

Pronunciation assessment in foreign language learning: Reliability and scoring bias in human–generative AI evaluation

Primary research

#1105

T1new
Topic
unassigned (set during synthesis)
First seen
2026-08-01 07:15:59
Last seen
2026-08-01 07:15:59

Source raw items (1)

  • Semantic Scholar2026-08-01 07:15:27
    Pronunciation assessment in foreign language learning: Reliability and scoring bias in human–generative AI evaluation

    This study examines the reliability and scoring bias of generative AI (Gen-AI)-based pronunciation assessment compared with human raters, addressing whether AI-generated scores can be trusted in real educational settings. Sixty students participated in a 12-week program. A total of 180 pronunciation samples were evaluated across eight subcomponents (individual phonemes, stress, rhythm, intonation, linking, reduction, fluency, and clarity) by three standardized human raters and Gen-AI using the same 7-point rubric. Quantitative analyses (intraclass correlation coefficients, paired-samples t-tests, and Pearson correlations) assessed reliability and bias, while semi-structured interviews with raters provided explanatory qualitative insights. Gen-AI demonstrated moderate reliability with human raters across most components, showing the highest agreement in fluency and the weakest in individual phonemes. However, Gen-AI consistently assigned significantly higher scores than human raters across all subcomponents. Qualitative findings revealed that discrepancies originated from Gen-AI’s limited discriminative power, systematic flaws in handling missing data, decontextualized scoring approach, lack of sensitivity to L1 interference, and inability to interpret pragmatic context. While Gen-AI cannot fully replace human expertise, it can function as a complementary tool for formative assessment and autonomous practice. A hybrid assessment model integrating Gen-AI’s efficiency with human raters’ contextual and interpretive insights is recommended for effective foreign language pronunciation education.