Fixed Deep Residual Networks with Multi-Objective Proximal Policy Optimization for Adaptive Trading – American Journal of Student Research

American Journal of Student Research

Fixed Deep Residual Networks with Multi-Objective Proximal Policy Optimization for Adaptive Trading

Publication Date : Aug-07-2026

DOI: 10.70251/HYJR2348.44727739


Author(s) :

Xichen Song.


Volume/Issue :
Volume 4
,
Issue 4
(Aug - 2026)



Abstract :

Adaptive trading systems must balance return generation with downside control while processing noisy, non stationary market data. This study evaluates a preference conditioned trading framework that combines a randomly initialized, permanently frozen deep residual network (DRN) with a single multi-objective proximal policy optimization (PPO) agent. Daily Bitcoin data from 2021 through 2025 were transformed into nine technical indicators and organized into 14 day windows. The frozen DRN mapped each window to a 32 dimensional representation, which was concatenated with profit and downside preference weights before entering the PPO actor critic policy. Four analyses were performed. A controlled seed 67 encoder ablation compared raw indicators, a fixed linear projection, a fixed non residual convolutional network, the fixed DRN, and a jointly trained DRN. The fixed DRN achieved the highest mean net profit across 20 preference evaluations (26.15%), whereas raw indicators produced the shallowest mean maximum drawdown (-5.15% versus -11.61% for the fixed DRN). Reward diagnostics showed that the directional profit and downside return components were of the same order of magnitude, with mean absolute values of 0.02194 and 0.01097, respectively. An ordered preference analysis showed measurable changes in policy probabilities and action composition, but deterministic actions were piecewise constant and maximum drawdown was unchanged across the seed 67 grid. In a common date seed 67 benchmark, the unified model produced 41.60% net profit, compared with 16.68% for a 20/50 day moving average strategy, 2.80% for the traditional fuzzy ensemble, and -0.06% for Buy-and- Hold. However, seven independently retrained seeds produced substantial uncertainty. At the neutral preference, the paired net profit difference favored the unified model by 7.73 percentage points, but the 95% bootstrap confidence interval included zero (-4.78 to 21.07), and the unified model exhibited deeper drawdowns on average. The results show that a fixed residual random projection can support profitable preference conditioned policies in some Bitcoin backtests, but they do not establish statistically robust superiority, improved maximum drawdown control, or a smooth continuous risk-return frontier.