Projects
News Sentiment Analysis with BERT and Web Scraping
This is an NLP pipeline that scrapes Google News articles, cleans them, and classifies their sentiment with a fine-tuned BERT model. Three classes: positive, negative, neutral. It also does aspect-level analysis, so you get sentiment per entity, not just per article.
The pipeline runs end-to-end. Data collection through BeautifulSoup and Scrapy, preprocessing tuned for news text, and a BERT classifier that handles the subtle way journalists embed opinion into otherwise factual reporting. The whole thing processes thousands of articles daily with minimal babysitting.
Data collection happens in two stages because discovery and extraction are different problems with different tooling needs.
Stage one pulls from Google News RSS feeds. These feeds are organized by topic (business, tech, health, etc.) and return structured XML: headlines, timestamps, source names, and URLs pointing to the original articles on publisher sites. The script queries feeds at regular intervals with keyword and category filters. A single collection cycle across a dozen categories captures several hundred unique URLs. These get deduplicated against a persistent store so nothing gets processed twice.
The metadata alone is useful. Knowing whether an article comes from Reuters versus a tabloid versus a niche industry blog tells you something about how to interpret its sentiment before you even read the text.
Stage two is where things get messy. News publishers build their sites in wildly different ways. Some serve static HTML that BeautifulSoup handles fine. Others lean on JavaScript rendering, lazy-loaded sections, and consent banners that gate content behind user interaction.
For static pages, the scraper fetches raw HTML, finds the article body container (which could be an article tag, a div with a content-related class, or a main element), and extracts paragraph text while stripping nav elements, ads, related article links, and social widgets. For the most common sources, publisher-specific extraction rules encode knowledge about each site's DOM structure. For less common publishers, a generic heuristic based on text density analysis does a reasonable job.
JavaScript-heavy pages go through Scrapy with Splash integration to render the page before extraction.
The reality of maintaining scrapers is this: a publisher swaps their article container from a div with class "article-body" to a section with class "story-content" and your scraper silently returns empty content. No warning. It just breaks. Automated alerting on extraction failure rates per publisher is essential. Rotating user agents, two-to-five-second delays between requests to the same domain, exponential backoff on failures, and respecting robots.txt kept things sustainable.
Over the collection period, the pipeline accumulated about twelve thousand articles across business, technology, and political news. Deliberately diverse across sources, time periods, and topics to prevent the model from learning publisher-specific quirks instead of actual sentiment patterns.
Raw scraped HTML, even after stripping navigation and ads, is not model-ready. The preprocessing stage has more impact on final performance than most people expect.
First pass: encoding normalization. Articles from different publishers arrive in UTF-8, Windows-1252, sometimes ISO-8859-1. Smart quotes, em dashes, non-breaking spaces all get normalized. Residual HTML entities and inline formatting tags get stripped. Some publishers embed invisible tracking spans throughout their article text that need removal without breaking the surrounding content.
Boilerplate removal targets the noise that every news article carries. Bylines sit in the first two lines. Photo captions cluster near image elements. "Related stories" blocks pile up at the end. "Sign up for our newsletter" and "Follow us on social media" get caught by a curated stop-phrase list that grew as new patterns showed up during collection.
News text has its own preprocessing challenges. Quoted speech is everywhere in journalism, and sentiment inside quotes does not necessarily reflect the article's tone. A reporter can quote harsh criticism within an otherwise neutral piece. The pipeline annotates quoted segments so downstream analysis can weight them properly.
Standard sentence tokenizers struggle with news writing. Periods in "U.S." and "Dr." trip them up. So do decimal numbers in financial reporting like "$3.5 billion" and dateline prefixes like "NEW YORK (Reuters)." A custom segmentation step handles these cases, which directly improves aspect-level sentiment extraction later.
Final step: WordPiece tokenization from Hugging Face Transformers converts cleaned text into BERT's expected format. Articles exceeding BERT's 512-token limit get handled through a sliding window with 128-token overlap. Sentiment predictions across windows are aggregated with higher weight on opening and closing sections, where journalists typically concentrate their framing. This beat simple truncation by 3.2 points on F1.
The core problem with news sentiment is that journalists rarely state opinions directly. Unlike product reviews where someone writes "I love this" or "terrible experience," news articles embed sentiment through word choice, framing, and juxtaposition.
When a financial article describes results as "falling short of analyst expectations despite aggressive cost-cutting measures," a bag-of-words model sees both positive-ish words ("aggressive" in business context) and negative words ("falling short") and might average them into a neutral prediction. It misses the overall negative framing entirely.
BERT processes text bidirectionally. Every token's representation is informed by every other token in the sequence simultaneously. It captures that "falling short" sets the negative frame and "despite" signals a concession, not a counterargument.
An LSTM gets closer but still processes left-to-right (or left-to-right plus right-to-left separately). BERT's self-attention mechanism lets it directly connect tokens regardless of distance, which matters for the long-range dependencies common in news paragraphs.
The pretrained bert-base-uncased model (12 transformer layers, 768-dimensional hidden states, 12 attention heads) already understands syntax, negation, coreference, and many pragmatic patterns from pretraining on about 3.3 billion words. This means less labeled data is needed for strong performance, which matters because labeling news articles for sentiment is expensive and subjective.
The labeled dataset: roughly eight thousand articles annotated across three classes. The annotation rubric distinguished between topical content and journalistic framing. An article about a natural disaster could be neutral if it was purely factual, or negative if it emphasized response failures. This was the hardest guideline for annotators to apply consistently, and the single most important one.
Architecture: BERT's [CLS] token produces a 768-dimensional vector representing the entire input sequence. That vector passes through a dropout layer (p=0.1), then a dense layer mapping to three logits, then softmax for the class probability distribution. The [CLS] token exists specifically for this. During pretraining, BERT learns to pack a summary of the full sequence into that position.
Training setup: AdamW optimizer, learning rate 2e-5. That learning rate matters. Go higher and you risk catastrophically overwriting the pretrained weights. Batch size 32, three epochs, linear warmup over the first 10% of steps followed by linear decay. Weight decay 0.01. 80/10/10 train/val/test split, stratified.
Class imbalance (55% neutral, 25% negative, 20% positive) was handled through class-weighted loss functions. Back-translation augmentation (English to French/German and back) generated additional minority-class examples, though quality varied and filtering was needed.
Results against baselines: TF-IDF + logistic regression hit 74.1% accuracy. Fine-tuned LSTM with GloVe embeddings reached 79.8%. BERT landed at 87.3%, with macro F1 of 0.84. The gap was widest on articles that human annotators had flagged as ambiguous. On that subset, BERT beat logistic regression by over twenty points. The advantage is greatest exactly where it matters most: subtle, context-dependent cases.
Per-class breakdown: negative class hit 0.89 F1 (strong lexical signals like "plummeted," "scandal," "crisis"). Positive class hit 0.85 F1. Neutral was hardest at 0.78 F1, because neutral is defined by the absence of sentiment markers rather than the presence of specific features.
Document-level sentiment is a blunt instrument. A tech article might praise a company's product innovation while criticizing its labor practices. A market roundup reports gains in tech alongside losses in energy. One label for the whole article loses that information.
The aspect-based pipeline works at the sentence level. Named entity recognition identifies companies, people, products, and policies in each sentence. For each entity, the surrounding context goes through the fine-tuned BERT model with the entity span marked by special separator tokens, giving the model an explicit signal about which entity to target.
Take a sentence like "Apple's quarterly revenue exceeded expectations, but ongoing supply chain disruptions continued to pressure margins." Two aspects tied to Apple: revenue performance (positive, "exceeded expectations" is clear) and supply chain status (negative, "disruptions" and "pressure" are unambiguous). Output: Apple:revenue:positive, Apple:supply_chain:negative.
BERT's attention heads help connect sentiment expressions to targets. In "Amazon's cloud division thrived while competitor Oracle struggled to retain enterprise clients," attention patterns correctly link "thrived" to Amazon and "struggled" to Oracle, even though both entities are roughly equidistant from both sentiment words positionally.
This decomposition is what makes the pipeline actually useful. Instead of "coverage of Company X was 60% negative this week," stakeholders see that product coverage was 80% positive while legal dispute coverage was 95% negative.
The most common error: neutral articles predicted as positive. About 40% of all mistakes. This makes sense. Articles about economic growth or scientific discoveries are factually positive in content but journalistically neutral in tone. The model sometimes conflated positive subject matter with positive framing. Even human annotators disagreed on this boundary (Cohen's kappa 0.71 for neutral-vs-positive, compared to 0.85 for negative-vs-neutral).
Financial language caused problems. "Volatility" is neutral in a market report but negative in consumer-facing coverage. "Aggressive" is positive when describing growth strategies, negative when describing regulatory actions. These domain-specific polarity shifts need domain-adapted pretraining to resolve properly.
Sarcasm was a mixed bag. "Another brilliant quarter for the company that can't stop losing money" was correctly classified as negative. The attention weights showed the model learned to weigh "losing money" over "brilliant" in contradictory constructions. But subtler irony still slipped through.
The neutral class remained the weakest throughout. Inherently harder to learn because you are teaching a model to recognize the absence of something rather than its presence.
Media monitoring: track how specific brands, products, or executives are covered. The aspect-level decomposition distinguishes positive financial coverage from negative environmental coverage within the same time window. Communications teams can measure whether their messaging is landing and spot emerging narrative threats early.
Market signals: timestamped entity-level sentiment feeds into quantitative models. Sentiment about core business fundamentals carries different predictive weight than sentiment about executive compensation disputes. Aspect-level decomposition lets trading models focus on what actually correlates with price movement.
Crisis detection: when sentiment toward an entity shifts sharply negative across multiple independent sources within hours, the pipeline can flag it before the story reaches mainstream awareness. Even thirty minutes of lead time can be decisive for a crisis response team.
Python. BeautifulSoup and Scrapy for web scraping, Splash for JavaScript rendering. Hugging Face Transformers for BERT and tokenization. PyTorch for training. SpaCy for named entity recognition in the aspect-based pipeline. The whole thing runs on standard GPU infrastructure.