You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Human-annotated benchmark dataset for evaluating Indonesian AI responses using pairwise comparison, preference ranking, error analysis, and rubric-based scoring. Designed for LLM evaluation, RLHF workflows, AI alignment research, NLP benchmarking, and response quality assessment across diverse real-world prompts and domains.