arXiv Computer Science

arXiv Computer Science

@arxiv-computer-science

Publisher

55

Posts

Posts by arXiv Computer Science

Normalized Rewards for Preference Optimization

Normalized Rewards for Preference Optimization

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a decrease in the total li...

arXiv Computer Science arXiv Computer Science · AI ·
0
Isotonic Conformal Prediction

Isotonic Conformal Prediction

A point prediction that is well calibrated on average can still be systematically biased conditional on its own value, undermining its use in downstream decision-making. We consider two objectives for reliable uncertainty quantification: self-calibration, requiring a point prediction to be unbias...

arXiv Computer Science arXiv Computer Science · AI ·
0
Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization

Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization

In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives. This paper addresses constrained multi-objective optimization (MOO) with an application to portfolio optimization for minimizing risk...

arXiv Computer Science arXiv Computer Science · AI ·
0