<?xml version="1.0" encoding="UTF-8"?>
		<www.wjpsonline.org>
		<Title>Reinforcement Learning-Based Self-Improving LLM Agents for Autonomous Task Optimization</Title>
		<Author>Ravi Kumar Kottala</Author>
		<Volume>2</Volume>
		<Issue>4 ( October - December )</Issue>
		<Abstract>The previous accepted convention of largo Language Model LLM agents used for autonomous task execution has been that they are static systems their behaviour when used at inference time is fixed and does not depend on feedback signals they receive during run time The paper introduces a reinforcement learning RL approach to designing a selfimproving LLM agent framework that allows agents to selectively improve action selection policies rather than finetuning LLM parameters by accepting scalar rewards based on the success of their action The proposed architecture combines a GPT4LLaMA task planner with an episodic memory module that looks up previously seen stateaction pairs and an actorcritic policy network trained using proximal policy optimisation PPO to concurrently optimise for task accuracy execution efficiency and operational cost represented by a composite reward function Rsa  A  E  C The framework is tested against a baseline static agent and a memoryaugmented agent against three benchmarks HumanEval GSM8K and HotpotQA Experimental results show the RL agent attains 97 task accuracy 95 learning efficiency and 1650 ms prediction latency on average due to which 15 high and 20 high improvements are achieved over the standard agent and prediction latency stays in acceptable range for the conversations when compared with the humans Learning convergence is reached after 1012 training episodes while the policy updater of PPO is stable along the training trajectory due to its KLdivergence KLD clipping mechanism This paper demonstrates that reinforcement learning is a tractable method for the continual autonomous optimization of deployed LLM agents</Abstract>
		<permissions>
<copyright-statement>Copyright (c) World Journal of Pharmaceutical Seiences. All rights reserved</copyright-statement>
<copyright-year>2026</copyright-year>
</permissions>
		</www.wjpsonline.org>
		