Ruprecht-Karls-Universität Heidelberg
Bilder vom Neuenheimer Feld, Heidelberg und der Universität Heidelberg

Process Reward Modeling in LLMs

Module Description

Course Module Abbreviation Credit Points
BA-2010[100%|75%] CS-CL 6 LP
BA-2010[50%] BS-CL 6 LP
BA-2010[25%] BS-AC 4 LP
BA-2010 AS-CL 8 LP
Master SS-CL-TAC 8 LP
Lecturer Lei Tang
Module Type Proseminar / Hauptseminar
Language English
First Session 16.04.2026
Time and Place Thursday, 15:15 - 16:45,
INF 326 / SR 27
Commitment Period tbd.

Participants

All advanced CL Bachelor students and all CL master students. Students from MSc Data and Computer Science or MSc Scientific Computing with Field of Application Computational Linguistics are welcome after getting permission from the lecturer. MSc Scientific Computing students can only take the course as HS for 8 LP.  If the seminar should be oversubscribed, CL students will have priority.  

Prerequisites for Participation

  • Statistical Natural Language Processing
  • Basic Knowledge in Neural Networks

Assessment

  • Presentation (50%)
  • Project (50%)

Content

Recent advances in Large Language Models (LLMs) suggest that Process Reward Models (PRMs) offer a promising approach to verifying intermediate reasoning steps and enhancing model performance. In this seminar, we will provide a systematic review of foundational and influential papers on PRMs.

Schedule


intro_slide

» More Materials

zum Seitenanfang


© Copyright Universität Heidelberg | Impressum | Webmaster

This page was last modified on Saturday July 11, 2026
Date Material Presenter
16.4 Introduction and Organizational Information Lei Tang
23.4. No Session. Prepare the presentation.
30.4. PRMs W/o Human Annotations
[1] Hunteraker et al., 2023. Let’s verify step by step.
[2] Wang et al., ACL 2024. Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Gang Yan,Xinyao Peng
Bingyue Li, Qingyang Cao, Siqing Cai
07.5. Generative PRM with Code
[1] Li et al., 2025. CodePRM: ExecutionFeedback-enhanced Process Reward Model for Code Generation.
[2] Zhao et al., GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

Shiya Feng, Zhenglin Lin, Yiting Tong
Jiaze Li, Xiaohui Wang, Yuyan Zhu
21.5. Step-wise DPO, and Incorporating PRM with DPO
[1] Lai et al., 2025. Step-DPO: Step-wise preference optimization for long-chain reasoning of LLMs.
[2] She et al., EMNLP 2025. R-PRM: Reasoning-Driven Process Reward Modeling.
Mohammed Arsh, Alejandro Petry Pacheco, Polina Kuznetcova
Yuzhou Shi, Wenshuang Hu, Xinyuan Ren
28.5. Incorporating PRM with Q-value, and Multilingual PRMs
[1] Li et al., Process Reward Model with Q-value Rankings.
[2] Wang., EMNLP findings 2025. Demystifying Multilingual Reasoning in Process Reward Modeling.
Keyan Chen, Zhikai Zhang
Georg Piersig, Amirreza Tarabkhah
11.6. Summary of current Progress, and a new PRM Bench
[1] Zhang et al., Lessons of Developing PRMs.
[2] Song et al., PRMBENCH: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.
Tianchen Wan
Binheng Zheng, Yuefeiyang Li
18.6. PRMs with RAG and Reward Tree
[1] Zhu et al., ACL findings 2025. Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning.
[2] Yin et al., ACL 2025. Dynamic and Generalizable Process Reward Modeling.
Pratik Goyal, Ashhad Raza Quadri
Ertugrul Taparci, Yahor Lahunovich
25.6. Error-aware PRM
[1] Tej Deep Pala et al., EMNLP findings 2025. Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision.
[2] Yang et al., EMNLP findings 2025. Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning.
Qin Yan, Ningyun Chen, Yiming Li
Anni Wang, Jiayi Ji, Tiange Lyu
02.7. Outcome-supervised Value Model, and Learning from ORM
[1] Yu., et al., OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.
[2] Xie., ACL 2025. From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment.
Mario Kuzmanov,
Yimin Yan, Zhiheng Lin
09.7. Process Labels from ORMs
[1] Lu et al., 2024. AutoPSV: Automated process-supervised verifier.
[2] Yuan et al., ICML 2025. Free Process Rewards without Process Labels
Yunhe Dong, Xiaoci Zhang
Haider Irfan
16.7. Project Preparation
15:15-15:30:Gang Yan,Xinyao Peng
15:30-15:45:Bingyue Li, Qingyang Cao, Siqing Cai
15:45-16:00:Shiya Feng, Zhenglin Lin, Yiting Tong
16:00-16:15:Qin Yan, Ningyun Chen, Yiming Li
16:15-16:30:Mohammed Arsh, Alejandro Petry Pacheco, Polina Kuznetcova
16:30-16:45:Yuzhou Shi, Wenshuang Hu, Xinyuan Ren
23.7. Project Preparation
15:15-15:30:Keyan Chen, Zhikai Zhang
15:30-15:45:Georg Piersig, Amirreza Tarabkhah
15:45-16:00:Tianchen Wan
16:00-16:15:Binheng Zheng, Yuefeiyang Li
16:15-16:30:Pratik Goyal, Ashhad Raza Quadri
16:30-16:45:Ertugrul Taparci, Yahor Lahunovich
30.7. Project Preparation
15:15-15:30:Jiaze Li, Xiaohui Wang, Yuyan Zhu
15:30-15:45:Anni Wang, Jiayi Ji, Tiange Lyu
15:45-16:00:Mario Kuzmanov
16:00-16:15:Yimin Yan, Zhiheng Lin
16:15-16:30:Yunhe Dong, Xiaoci Zhang
16:30-16:45:Haider Irfan