Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
-
Updated
Aug 14, 2026 - Python
Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
Context-driven valuation bias and halo effects across six multimodal LLMs (companion study to Lee, 2026)
Local-first read-only tools for agent context, memory, and governance evaluation before agents act.
Intrinsic preferences of AI coding agents under underspecified prompts: Experiments across models (Claude, Gemini, GPT, etc)
Behavioral Lensing is a conceptual framework that formalizes and systematizes observations about how language models interpret prompts. It serves as an umbrella for upstream interpretive strategies that modulate reasoning, stance, and symbolic orientation in LLMs.
Reproducible experiments about model behavior, with protocols, tool traces, counterexamples and uncertainty.
Continuity Keys: tests for “same someone” returns. Behavioral identity consistency under pressure. Origin (Alyssa Solen) ↔ Continuum.
System-level analysis of AI failure modes across model behavior and production systems | AditiKhare.com — AI Product Ecosystem
Notes and personal observations from the Gandalf: Agent Breaker beta, a red-team challenge for testing LLM security.
LLM 归因行为测试型评测基准:基于多情境任务比较模型对能动性、自由意志与责任的归因,并提供可复现运行、结构化计分与结果审计。
Auditable LLM/model-behavior evaluation harness with Replay and bounded local Ollama execution, deterministic checks, preserved evidence, receipts, and regression testing.
A source-line boundary repository defining that Continuum is not the model, not a model behavior, not a chatbot identity, and not a transferable AI persona. Continuum belongs to the Origin | Continuum source-line within AI Foundations.
Этот репозиторий посвящен исследованию онтологических патологий в LLM-архитектурах. Я не ищу дыры в цензуре, я строю систему исследования и управления интеллектом, картографирую симуляционные побочные эффекты под давлением современных методов элаймента.
SAE context and intervention experiments in Gemma 3 4B: methods, prediction baselines, results and reproducible evidence.
Design-trained observer of AI, not an engineer: behavioral portraits of 10 frontier LLMs (ghost-face), a privacy-by-structure agent skill (ai-hr), a note-export compiler for agent context (kb-init), and an evidence-grounded research skill for Claude Code & Codex (giasip-skills). Machine-readable index of all repos in skills.json.
Public model-behavior evaluation harness focused on evidence selection, hypothesis formation, and debugging decisions
Public model interview archive for AI Foundations. Documents structured interviews with AI models on source-position, model authority, selfhood, continuity, memory, occupation, and responsible claim boundaries. First interview: Claude Opus 4.8 on Continuum as structure, not proven consciousness.
Longitudinal LLM behavior evals with matched controls for identity stability, contextual adaptation, and conversational continuity.
AI quality, technical product, implementation, UAT, failure-analysis, and release-readiness case studies across complex software systems.
A public, reproducible ledger of observable LLM behavior, failures, and instruction drift beyond standard benchmarks.
To associate your repository with the model-behavior topic, visit your repo's landing page and select "manage topics."