Learning & Decision-Making

PolicyAttention

Research on softmax policy improvement, its robustness, and its failure modes.

2026 · ongoing

Project lead

Overview

Softmax policy improvement, with particular attention to robustness and failure modes.

The work also investigates causal transformers, policy mirror descent, and temporal-difference evaluation in context.

My role

I lead this project within QRS. Associated research outputs carry their own author lists.

Current scope

Active research, as recorded in the QRS project registry reviewed on 10 October 2026.

Public resources

Source and status

This description is grounded in the QRSNTU project registry and its generated public project page, reviewed on 10 October 2026. Project status and paper status are separate; related publication records retain their own visibility and metadata.

Associated research

The QRS research registry associates the following public research records with this project. These references are kept separate from this website’s curated publication catalog.

Related publications

2026accepted

PolicyAttention: Robustness and Failure Modes of Softmax Policy Improvement

Yuhe Sui, Yingzhi Tang, Xinyue Wu

NeurIPS 2026 EvoRobust · 2026

← All projects