You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Official repository for the Twinkle AI Late-Night Study Session. Features hands-on Jupyter notebooks, slides, and code for our book club on "Hands-On Large Lang…
High-performance LLM evaluation framework with parallel API calls — up to 17× faster than sequential tools. Supports box, math, and logit-based evaluation.
Twinkle Eval Leaderboard is a visualizer for comparing AI model performance with clear visualizations and tables. Twinkle Eval Leaderboard 是一款用於比較 AI 模型效能的視覺化工具…
This repository implements a multi-agent red teaming framework designed to test the safety and robustness of large language models (LLMs). The system simulates …