Learning desk

MultiAgent EDU StackGather good sources. Teach what matters.
T5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better resultsT5TauricResearch/TradingAgentsT5A Man Who Invented Modern AI (Before Everyone Else) – Jürgen Schmidhuber [video]T5GPT-4 finished training four years ago todayT5AI Settles a 25 Year-Old Problem We Left BehindT5What it was like working on LLMs and security at Meta (2022-2026)T5Ask HN: How do you go from writing code to deploying with agents?T5What Happened: OpenAI and HuggingFaceT5Apple says Mac users in China can connect to Alibaba's Qwen AI serviceT5Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on SonnetT5The AI Apocalypse Is HereT3Auto mode is now the default in Claude Code for Pro, Max, and Team plansT5Show HN: Tura – Build agent that uses 80% less token and delivers better results
← Dispatches

Comparative Effectiveness of AI Tools in Enhancing Feedback Quality and Efficiency in Undergraduate Mathematics Assessments

Primary research

#886

T1new
Topic
unassigned (set during synthesis)
First seen
2026-07-28 07:16:45
Last seen
2026-07-28 07:16:45

Source raw items (1)

  • Semantic Scholar2026-07-28 07:16:04
    Comparative Effectiveness of AI Tools in Enhancing Feedback Quality and Efficiency in Undergraduate Mathematics Assessments

    High-quality, timely feedback is essential for students’ learning in mathematics; yet large class sizes and heavy workloads for teachers often result in students receiving brief, generic comments that fail to address misconceptions effectively. While artificial intelligence tools show promise for improving grading efficiency and feedback quality in STEM education, empirical evidence specific to undergraduate mathematics remains limited. This study conducted a rigorous comparative evaluation of two AI platforms — Gradescope (AI answer grouping and batch feedback) and StarGrader (generative AI with custom prompts for personalised feedback) — against traditional manual grading. The study analysed 6,922 student-question responses from three core undergraduate courses in Multivariable Calculus, Linear Algebra, and Differential Equations. Results demonstrated exceptionally high inter-rater reliability. Gradescope achieved 99.91% exact agreement with manual grading and 99.95% of scores within ±5 marks, while StarGrader attained 95.57% exact agreement and 96.44% practical reliability. Time savings were modest (average for Gradescope 22.2% and StarGrader 19.8%), but reached 30% for another course in Mathematical Analysis, and up to 35.1% in computation-heavy worksheets with refined prompts. Prompt engineering turned out to be critical especially for StarGrader, with accuracy improving from 43.75% to 94.17% through systematic refinement. Both student surveys and instructor ratings indicated more positive perceptions of AI-generated feedback regarding clarity, personalisation, and usefulness. The findings indicate that AI tools, when supported by effective prompt strategies and human oversight, can substantially enhance feedback quality and efficiency in mathematics assessments. Two evidence-based user manuals were developed to support wider adoption of the tools.