NLP · Applied AI · Data Engineering · Full-stack
Stock Forum Summarization & Sentiment Analysis System
A multi-agent LLM pipeline that turns noisy forum chatter into structured stock signals
Role
Fully designed and built by me — collection, data pipeline, agents, APIs, deployment and UI.
When
Personal project · used by an internal team
Tech
- Python
- LangChain
- FastAPI
- SQLAlchemy / SQLModel
- SQLite
- EasyOCR
- Transformers
- StockBERT
- Next.js
- Tailwind CSS
- Docker
- Nginx
- Certbot
An end-to-end system built independently: browser-style collection from a private stock forum, parsing / deduplication / OCR, a SQLite store, a LangChain multi-agent analysis pipeline with finance-domain sentiment, and a dashboard for reviewing and exporting results. It ran for an internal team rather than staying a local prototype.
Numbers that are real
100,000+
Historical messages collected
several thousand / day
Ongoing collection
internal team
Used by
What I did
- Designed a LangChain multi-agent pipeline: a low-cost filter agent removes off-topic chat, a second agent checks whether a company name is really being used as a stock reference, and an analysis agent returns structured output with the stock, the reasoning and the sentiment.
- Found experimentally that asking the analysis agent for a written rationale alongside its conclusion produced more reliable outputs; structured rationale prompting is now part of the system design.
- Built the collection and cleaning pipeline: browser-style HTTP requests with randomised request intervals, plus parsing, deduplication, OCR and text extraction across messages, images and linked content.
- Stored 100,000+ historical messages in SQLite with ongoing collection of several thousand messages per day, exposed through backend APIs and scheduled jobs.
- Containerised and deployed the whole system on a lightweight server behind Nginx with Certbot TLS, including a dashboard to monitor runs, review outputs and export results.
Honest limitations
- Stock-entity attribution is the hard part: when one sentence mentions several companies, the model can attach an opinion to the wrong stock.
- Outputs were reviewed by financially experienced team members against subsequent market movement. The top bullish calls did not consistently rise, so this is an information and prioritisation signal — not a prediction, and not an accuracy claim.
- Next step: replace the two low-cost LLM filter agents with lighter dedicated NLP classifiers, and add entity linking / target-dependent sentiment for multi-entity sentences.