Joongkyu Lee
Joongkyu Lee
Home
Publications & Preprints
Light
Dark
Automatic
1
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. …
Deokgyu Yoon
,
Hyungkyu Kang
,
Joongkyu Lee
,
Byeongchan Kim
,
Gyungin Shin
,
Sungrae Park
,
Min-hwan Oh
PDF
Code
Block-Sphere Vector Quantization
Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast …
Heesang Ann
,
Joongkyu Lee
,
Min-hwan Oh
PDF
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
We study optimal experimental design for multinomial logit (MNL) bandits, where an agent repeatedly selects a subset of $K$ items from …
Joongkyu Lee
,
Min-hwan Oh
PDF
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of …
Joongkyu Lee
,
Seouh-won Yi
,
Min-hwan Oh
PDF
True Impact of Cascade Length in Contextual Cascading Bandits
We revisit the contextual cascading bandit, where a learning agent recommends an ordered list ($\text{\textit{cascade}}$) of items and …
Hyunjun Choi
,
Joongkyu Lee
,
Min-hwan Oh
PDF
Combinatorial Reinforcement Learning with Preference Feedback
In this paper, we consider combinatorial reinforcement learning with preference feedback,where a learning agent sequentially offers an …
Joongkyu Lee
,
Min-hwan Oh
PDF
Improved Online Confidence Bounds for Multinomial Logistic Bandits
In this paper, we propose an improved online confidence bound for multinomial logistic (MNL) models and apply this result to MNL …
Joongkyu Lee
,
Min-hwan Oh
PDF
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit (Top 0.2%, 32/15671)
In this paper, we study the contextual multinomial logit (MNL) bandit problem in which a learning agent sequentially selects an …
Joongkyu Lee
,
Min-hwan Oh
PDF
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
We study reinforcement learning with
multinomial logistic
(MNL) function approximation where the underlying transition probability …
Wooseong Cho
,
Taehyun Hwang
,
Joongkyu Lee
,
Min-hwan Oh
PDF
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
In this paper, we first challenge the common premise that linear MDPs always induce performance guarantees independent of the state …
Joongkyu Lee
,
Min-hwan Oh
PDF
»
Cite
×