What RedViz Is
Why We Built It
The Data
What the attack corpus looks likeThe CohereLabs AYA red-teaming corpus: 7,419 human-written adversarial prompts across 8 languages and 9 harm categories. Prompts carry more than one harm label, so the category counts sum beyond the 7,419 total, while the language counts partition it exactly.
Prompts per harm category (multi-label)
Four Dashboard Modules
What We Found
The Multilingual Angle
Tech Stack
The Team
- Anuj Gupta (anujg2): Lead developer and system architect
- Rohini Das (rohinida): Data analysis and visualization design
- Iskander Sergazin (isergazi): Model inference pipeline and interpretability modules
- Sihan He (sihanhe): Dataset curation and multilingual analysis
References
- The AYA red-teaming corpus and its global-versus-local harm annotations come from Aakanksha et al., "The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm," EMNLP 2024 (arXiv:2406.18682); the data is on Hugging Face.
- JailbreakBench: Chao et al., "JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models," NeurIPS 2024 (jailbreakbench.github.io).
- HarmBench: Mazeika et al., "HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal," 2024 (arXiv:2402.04249).
- Chu et al., "Comprehensive Assessment of Jailbreak Attacks Against LLMs," 2024.
- Belaire, Sinha, and Varakantham, "Automatic LLM Red Teaming," 2025 (arXiv:2508.04451).
- Generation runs on TinyLlama-1.1B-Chat and Llama; toxicity is scored with unitary/toxic-bert. Built with Streamlit and Hugging Face Transformers. Corpus figures above are from the team's own exploratory analysis of the AYA dataset.