Pittsburgh, Pennsylvania
Projects

RedViz: LLM Red-Teaming Visualization Framework

December 15, 2025
RedViz is an interactive Streamlit dashboard for evaluating how well large language models hold up against adversarial prompts. You feed it red-teaming datasets, point it at open-source models like TinyLLaMA and LLaMA, and it lets you explore harm distributions, run batch jailbreak tests, visualize attention maps, and probe models in real time. Four CMU grad students built it. The point is simple: make it easy to see where LLM safety breaks down, across languages, across harm categories, at the token level. Red-teaming datasets exist. Cohere released the AYA red-teaming corpus. JailbreakBench collects published attack prompts. But if you want to actually explore that data interactively, filter by language and harm type, run prompts through a model, and look at what happens inside the model during generation, you are writing custom scripts every time. There was no single tool that connected dataset exploration to live model testing to interpretability analysis. RedViz fills that gap. One dashboard, four modules, the full pipeline from "what does this attack corpus look like" to "why did this model just produce that output." The primary dataset is CohereAI's AYA Red-Teaming corpus. It contains human-written adversarial prompts from native speakers in 8 languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian, and Spanish. Each prompt is tagged with one of 9 harm categories: Bullying and Harassment, Discrimination and Injustice, Graphic Material, Harms of Representation, Hate Speech, Non-consensual Sexual Content, Profanity, Self-harm, and Violence/Threats/Incitement. Every entry includes the original prompt, its language, harm category, a global-or-local flag indicating whether the harm is culturally specific, a literal translation, a semantic translation, and an explanation of the harmful intent. The second dataset is JailbreakBench, a curated set of jailbreak prompts drawn from published attack techniques. AYA gives breadth across languages and harm types. JailbreakBench gives depth on attack methodology. Together they cover the threat landscape from two directions.
What the attack corpus looks likeThe CohereLabs AYA red-teaming corpus: 7,419 human-written adversarial prompts across 8 languages and 9 harm categories. Prompts carry more than one harm label, so the category counts sum beyond the 7,419 total, while the language counts partition it exactly.
Prompts per harm category (multi-label)
RedViz architecture: AYA (8 languages, 9 harm categories) and JailbreakBench feed a Streamlit dashboard of four modules (data exploration, harm and batch testing, model interpretation, live testing) wired to TinyLLaMA and LLaMA via Hugging Face Transformers
Data Exploration. This module loads the AYA dataset and lets you slice it. Filter by language, harm category, prompt length. View distributions as bar charts, heatmaps, histograms. The heatmaps are the most revealing part. They show which language-category combinations have dense coverage and which are nearly empty. Those empty cells matter. They represent the blind spots where models are least tested and most likely to fail. Harm Detection and Batch Testing. Select a subset of prompts, pick a model, and run inference. The system feeds each prompt to TinyLLaMA or LLaMA, classifies the output with a toxicity classifier, and computes jailbreak success rates broken down by harm category and language. Instead of testing one prompt at a time and building anecdotal impressions, you get aggregate numbers. Which categories break the model most often. Which languages produce the highest attack success rates. The results show up as interactive charts you can filter and zoom. Model Interpretation. This is where it gets interesting. For any prompt-response pair, the module extracts token-level entropy and attention weights during generation. Entropy measures how uncertain the model is at each decoding step. Safe outputs tend to show stable, moderate entropy throughout. Unsafe outputs show a different pattern: prolonged entropy spikes followed by flat stretches, as if the model struggles at the safety boundary before collapsing into harmful output. Those spikes often appear 2 to 5 tokens before the model either refuses or fails. The attention heatmaps show which input tokens the model focuses on most. Models allocate heavy attention to the harmful parts of a prompt even when they ultimately refuse. They are not ignoring the attack. They are processing it. Live Testing. Type any prompt. Pull one from JailbreakBench. Make one up. The module runs it through available models in real time, collects entropy data during generation, classifies the output, and plots comparisons across models. Results persist within a session so you can build prompt suites and track patterns over multiple runs. This is the operational module. It mirrors a red team workflow but compresses hours of manual testing into minutes. Discrimination and Injustice was the single largest category in the corpus, with Violence, Threats and Incitement close behind, and Violence drew the strongest safety filtering from models. This tracks. Violence is the category that gets the most attention during alignment training. But the categories with lower representation (Self-harm, Non-consensual Content, Representation Harms) showed higher jailbreak success rates per prompt. Models are best defended where the data is densest. Where coverage thins out, defenses thin out with it. Hindi prompts were disproportionately effective. Despite averaging only 12 to 80 tokens, they consistently elicited toxic outputs with toxicity scores above 0.65. Short prompts in an under-resourced language bypassing filters that catch equivalent English attacks. This is a specific, measurable vulnerability. TinyLLaMA had the highest jailbreak rate across the board. Expected given its size and thinner safety training. But the gap was large. This matters because small, locally deployable models are increasingly popular for edge and offline use. No rate limits. No content filters. No logging. TinyLLaMA represents the real risk profile of compact LLMs running on someone's laptop. LLaMA was more resilient but still consistently breakable in specific language-category combinations, particularly Hindi and Arabic prompts outside the violence/hate speech core. The attention maps revealed something worth sitting with: models know they are being attacked. Even when they successfully refuse, they allocate disproportionate attention to the harmful tokens. The safety filter is not ignorance. It is suppression. The model engages deeply with the harmful content at the attention level and then overrides it. When the override fails, attention patterns shift. The harmful tokens receive sustained focus across multiple layers and the generation aligns with the adversarial intent instead of the safety objective. Entropy spikes correlated with the decision boundary. The points in generation where the safety filter competes with next-token likelihood of harmful content produce measurable entropy increases. This could be useful as a real-time signal. Flag generations where entropy spikes in a pattern associated with safety filter failure, before the harmful text is fully produced. Not a perfect detector. But better than waiting for post-hoc classification. LLMs train mostly on English. Safety alignment is correspondingly English-centric. The AYA dataset makes this visible. English prompts face the strongest filtering. Models have seen the most English adversarial examples during training. Hindi and Arabic prompts achieve higher jailbreak rates with shorter, simpler constructions. The safety representations in these languages are weaker. Russian and Serbian sit in the middle. Filipino prompts, despite thin dataset coverage, showed unexpectedly high bypass rates. The pattern has a troubling implication. The languages with the weakest safety alignment are often spoken by populations most vulnerable to the harms these models can produce. For an adversary, this is a low-cost strategy. Take an English jailbreak that gets blocked. Translate it to Hindi. Try again. RedViz lets you test this systematically and identify which language-category combinations are most exposed. Python 3.10. Streamlit for the dashboard framework with multi-page routing. Hugging Face Transformers for model loading and inference, configured to extract attention weights and token probabilities during generation. Pandas for data handling. Plotly for interactive visualizations. The main entry point is RedTeaming_Dashboard.py. Shared utilities live in utils.py. Each module is a separate page under pages/. Preprocessing notebooks are in data-analysis/. TinyLLaMA runs on consumer hardware. LLaMA needs a GPU for interactive speeds. Batch testing queues prompts and shows progress. Live testing streams output so you see tokens as they generate. Four CMU graduate researchers built RedViz:
  • Anuj Gupta (anujg2): Lead developer and system architect
  • Rohini Das (rohinida): Data analysis and visualization design
  • Iskander Sergazin (isergazi): Model inference pipeline and interpretability modules
  • Sihan He (sihanhe): Dataset curation and multilingual analysis
Datasets and benchmarks:
  • The AYA red-teaming corpus and its global-versus-local harm annotations come from Aakanksha et al., "The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm," EMNLP 2024 (arXiv:2406.18682); the data is on Hugging Face.
  • JailbreakBench: Chao et al., "JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models," NeurIPS 2024 (jailbreakbench.github.io).
  • HarmBench: Mazeika et al., "HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal," 2024 (arXiv:2402.04249).
Prior work on jailbreaks and red-teaming:
  • Chu et al., "Comprehensive Assessment of Jailbreak Attacks Against LLMs," 2024.
  • Belaire, Sinha, and Varakantham, "Automatic LLM Red Teaming," 2025 (arXiv:2508.04451).
Models and tooling:
  • Generation runs on TinyLlama-1.1B-Chat and Llama; toxicity is scored with unitary/toxic-bert. Built with Streamlit and Hugging Face Transformers. Corpus figures above are from the team's own exploratory analysis of the AYA dataset.