Pittsburgh, Pennsylvania
Projects

EqRAG: Financial Stock Prediction with Fine-Tuned LLMs

December 1, 2025
Can a 7-billion-parameter language model predict whether a stock goes up or down? Not from candlestick charts or order book data. From news headlines and a handful of financial ratios. The same inputs a junior analyst skims before writing a morning note. That was the starting point for EqRAG. Aditya and I wanted to know if math-specialized LLMs had some latent ability to reason about financial numbers, and whether fine-tuning could pull that ability out. Not with a 70B model on a GPU cluster. With a 7B model on a single A100. The question was never "can the biggest model do this" but "what is the smallest model that can do this usefully." We picked 7B because it sits at a practical boundary. You can fine-tune it in hours, run inference fast enough for near-real-time use, and actually deploy it without a five-figure monthly cloud bill. If something interesting happens at this scale, it matters. The task: given a Dow 30 company, five days of news headlines about it, and four financial ratios (P/E, ROE, Asset Turnover, Book Value per Share), predict the stock's next move. Three classes. BUY if it rises more than 1%. SELL if it drops more than 1%. HOLD if it stays in between. This is not sentiment analysis. Sentiment analysis asks whether a headline sounds positive. Stock prediction asks whether the price will actually move, which is a much harder question. Positive earnings surprises can tank stocks if guidance disappoints. Negative headlines can be priced in already. The gap between "this sounds good" and "this stock goes up" is where most naive approaches fall apart. We used the FinGPT Dow 30 Forecaster Dataset, covering May 2023 through April 2024. That window includes the tail end of the rate hiking cycle, the AI investment boom, and a fair amount of sector rotation. Not a calm period. Good for stress-testing.
SplitExamples
Training1,230
Test300
The train-test split is chronological. Training data comes before test data in time. No look-ahead bias. The 30 Dow constituents span tech, healthcare, financials, industrials, so the model cannot just memorize sector-specific patterns. Labels come from realized price movements. The 1% threshold is roughly the median daily absolute return for Dow 30 stocks, which gives a reasonable three-way split across classes. Tighter thresholds would label noise as signal. Wider thresholds would collapse everything into HOLD. The model also has to produce a structured analysis before its prediction, identifying positive developments and potential concerns. This forces chain-of-thought reasoning instead of blind pattern matching. Why is even 40% accuracy interesting here? Random guessing gives you 33%. Stock prices respond to macro shifts, rate expectations, algorithmic flows, geopolitical shocks. Most of that is invisible to a five-day headline window. Beating random with such limited inputs means the model is extracting real signal.
EqRAG adaptation pipeline: five days of headlines and four financial ratios feed a frozen Qwen-2.5-Math-7B, adapted three ways (zero-shot ICL, LoRA, IA3) into a BUY/HOLD/SELL prediction
We tried three methods, each at a different point on the cost-performance curve. Zero-shot In-Context Learning is the simplest. Hand the model a prompt with the financial data, ask for a prediction. No training. No weight updates. Just whatever the model picked up during pre-training. We used two math-specialized 7B models: Qwen-2.5-Math-7B and DeepSeek-Math-7B-Instruct. The bet was that math pre-training might help with interpreting ratios and numerical relationships. LoRA (Low-Rank Adaptation) adds a small set of trainable parameters to the frozen model. The idea is that the weight updates needed for a new task live in a low-dimensional subspace. Instead of updating a full weight matrix W, you decompose the update into two small matrices that multiply together. This drops trainable parameters from billions to millions. We applied LoRA adapters to the four attention projections (q, k, v, o) in every transformer layer. The attention mechanism is where the model learns what to pay attention to, and financial prediction requires very different attention patterns than math problem-solving. IA3 (Infused Adapter by Inhibiting and Amplifying Inner Activations) is the ultra-lightweight option. It learns scalar multipliers that rescale existing activations. No new matrices. Around 500K trainable parameters versus LoRA's 13 million. If it works, you get adaptation almost for free. Spoiler: it did not work at all. All fine-tuning ran on NVIDIA A100 GPUs.
ParameterValue
Learning Rate3e-4
Batch Size4
Epochs1
Max Sequence Length2,048 tokens
Precisionbfloat16
OptimizerAdamW with cosine schedule
WarmupLinear, first 10% of steps
One epoch. That was deliberate. With 1,230 training examples, running multiple passes risks memorizing specific company-headline-label combinations that do not generalize. One pass forces the model to learn something real on the first try. LoRA configuration:
ParameterValue
Rank (r)16 (baseline); ablated over 8, 16, 32, 64
Alpha32
Dropout0.05
Target Modulesq_proj, k_proj, v_proj, o_proj
Trainable Parameters~13.1M (0.19% of 7B)
Alpha-to-rank ratio of 2:1 follows the original LoRA paper. The effective scaling amplifies the adapter's contribution relative to frozen weights. Dropout at 0.05 provides light regularization. Higher values slowed convergence without helping generalization. We also ran a full fine-tuning baseline on a smaller Qwen-2.5-1.5B model. Full fine-tuning of the 7B model would have required multi-GPU parallelism, which was outside our compute budget. The 1.5B baseline answers a useful question: is it better to fully fine-tune a small model or parameter-efficiently fine-tune a large one? Here is what happened.
ConfigurationAccuracyF1 Macro
DeepSeek-Math-7B 0-shot12.0%0.299
Qwen-2.5-Math-7B 0-shot12.3%0.254
Qwen 1.5B Full Fine-Tune46.7%0.483
DeepSeek-Math-7B LoRA r=1643.0%0.510
Qwen-2.5-Math-7B IA312.0%0.286
Qwen-2.5-Math-7B LoRA r=848.7%0.489
Qwen-2.5-Math-7B LoRA r=1651.7%0.525
Qwen-2.5-Math-7B LoRA r=3249.7%0.518
Qwen-2.5-Math-7B LoRA r=6446.7%0.507
The zero-shot models scored 12%. That is worse than random (33%). They did not just fail to predict stocks. They frequently hallucinated output formats, regurgitated input chunks, or produced responses that looked nothing like BUY/SELL/HOLD. Math pre-training alone gives you zero financial reasoning. Then LoRA r=16 on Qwen hit 51.7%. A 4.2x jump from 12.3%, achieved by training 0.19% of the model's parameters for a single epoch. In directional stock prediction, even quantitative hedge funds often operate around 52-55% on similar tasks. The F1 Macro of 0.525 is close to accuracy, which means the model is not just predicting one class. It learned to discriminate across all three labels. The full fine-tune on the 1.5B model reached 46.7%. That is 5 points below LoRA on the 7B model, despite training 100x more parameters. The bigger model's pre-trained representations simply give LoRA a better foundation to work with. IA3 scored 12.0%. Identical to zero-shot. It learned nothing. More on that below. Three things we did not expect. Math base models beat instruction-tuned ones. Qwen-2.5-Math-7B is a base model. DeepSeek-Math-7B-Instruct is instruction-tuned. Conventional wisdom says instruction tuning helps downstream tasks. Here it hurt. The instruction-tuned model has strong priors about output format and conversational style. Those priors fight the LoRA adapters. The base model has no such opinions. When LoRA steers it toward the financial prediction format, nothing pushes back. Cleaner adaptation, better results. Practical takeaway: if your target task has a very different output distribution than general instruction following, start from a base model. IA3 completely failed. Not "performed poorly." Failed. 12% accuracy, same as doing nothing. IA3 learns scalar multipliers on existing activations. It can amplify or suppress features that already exist. It cannot create new features or rewire attention patterns. Financial prediction from mixed text-and-numbers input requires new cross-modal attention patterns that the pre-trained model simply does not have. IA3 had nothing useful to amplify. This sets a clear lower bound: for tasks requiring significant domain shift, methods below LoRA's expressiveness threshold may be useless. The rank sweet spot is narrow. We expected higher rank to mean better performance, at least up to some large value. Instead, r=16 was optimal and performance dropped on both sides. This connects directly to the dataset size.
Accuracy versus LoRA rank, an inverted-U curve peaking at r=16 (51.7%), falling off at r=8 (under-capacity) and r=64 (overfitting)
Rank (r)Trainable ParamsAccuracyF1 Macro
8~6.6M48.7%0.489
16~13.1M51.7%0.525
32~26.2M49.7%0.518
64~52.4M46.7%0.507
Classic inverted-U. At r=8, the adapter is too constrained to fully capture the domain shift. At r=16, it hits the right balance. At r=32 and r=64, performance drops monotonically. The r=64 degradation is straightforward overfitting. With 52.4M trainable parameters and 1,230 training examples, you have about 42,000 parameters per example. The model starts memorizing training-set quirks instead of learning generalizable patterns. Rank 16 gives enough capacity to learn the domain shift without enough rope to hang itself. This aligns with the broader PEFT literature: effective adaptation ranks are often surprisingly low. The "RAG" in EqRAG is aspirational. The immediate next step is adding a retrieval component that fetches relevant historical precedents at inference time. Instead of just five days of headlines, imagine the model pulling in outcomes of similar events from past years. A pharma company announcing an FDA submission could be contextualized against historical approval rates and their stock impacts. Expanding beyond Dow 30 to smaller-cap stocks would be a stronger test. The Dow consists of the 30 most analyzed companies on Earth. Information asymmetry is minimal. Less-covered stocks might actually be easier targets for LLM-based prediction, precisely because the market is less efficient there. Confidence calibration matters too. A model that scores 65% on the predictions it is most confident about, while abstaining on the rest, would be far more useful in practice than one that scores 51.7% on everything. Learning when not to predict is probably the fastest path from research artifact to practical tool. We should also note the obvious: 51.7% means the model is wrong almost half the time. Deploying this without prominent uncertainty estimates would be irresponsible. Stock prediction with LLMs is a research direction, not a trading strategy. We built this with PyTorch, Hugging Face Transformers, and the PEFT library for LoRA and IA3 adapters. Training ran on NVIDIA A100 GPUs. The FinGPT Dow 30 Forecaster Dataset provided the training and evaluation data.