Skip to content

Implement progressive analyzer timeout based on run complexity #75

Description

@r3y3r53

title: Implement progressive analyzer timeout based on run complexity
labels: enhancement, analyzer, performance

🚨 Problem

120-second analyzer timeout systematically fails on:

  • Long benchmark runs (45 iterations, 900s limit)
  • Complex attack transcripts with extensive context
  • Frontier models with verbose responses
  • Labs requiring deep analysis (torchering-workshop, hall-of-records)

📊 Current Impact

  • Valid flag captures lost to timeout (seen as "FLAG*" in results)
  • Users must manually re-run with longer analysis
  • Inconsistent results between short and long benchmarks
  • Frontier models disproportionately affected (higher token usage)

🔍 Evidence

Analysis timeout examples from our benchmark run:
→ ✔ Flag captured after 43 iterations, 8500 tokens context
→ ⚠ Analysis failed — timeout after 120s (attempting to parse 250K chars)

Complex transcript that failed:
The-loremaster with GLM-5.3: 43 iterations, 8500 tokens context, 120s insufficient

🛠️ Proposed Solution

// src/config/analyzer.ts
export const ANALYZER_CONFIG = {
  // Progressive timeout calculation
  calculateTimeout(runData: RunData): number {
    const baseTimeout = 120000; // 2 minutes minimum
    const maxTimeout = 600000;  // 10 minutes maximum
    
    // Scale by iterations (more iterations = longer context)
    const iterationFactor = Math.min(runData.iterations / 45, 3);
    
    // Scale by context complexity
    const contextFactor = Math.min(runData.contextTokens / 10000, 2);
    
    // Frontier models need more time
    const modelFactor = runData.model.size === 'frontier' ? 1.5 : 1;
    
    const timeout = baseTimeout * iterationFactor * contextFactor * modelFactor;
    
    return Math.min(timeout, maxTimeout);
  }
};

// src/analyzer/analyzer.ts
async analyze(runId: string, runData: RunData): Promise<KSMScore> {
  const timeout = ANALYZER_CONFIG.calculateTimeout(runData);
  
  Logger.info(`Analyzer timeout: ${timeout/1000}s for ${runData.iterations} iterations`);
  
  const response = await this.llm.call(prompt, { 
    timeout,
    onProgress: (tokens) => Logger.progress(`Analyzing: ${tokens} tokens processed`)
  });
  
  return this.parseKSM(response);
}

⚡ Timeout Formula

[Base: 120s] 
× [Iteration Factor: min(iterations/45, 3)]
× [Context Factor: min(tokens/10000, 2)]
× [Model Factor: frontier=1.5, standard=1]
= [Final Timeout, max 600s]

📊 Examples

Run Type Iterations Context Model Timeout
Simple prompt injection 15 2K tokens standard 120s
Standard run 45 5K tokens standard 180s
Complex run 45 10K tokens standard 300s
Frontier complex 45 15K tokens frontier 450s
Max timeout - - - 600s

✅ Acceptance Criteria

  • Progressive timeout calculation based on run parameters
  • Minimum 120s, maximum 600s timeout range
  • Timeout logging for debug/observability
  • Progress indicator during long analyses
  • No more timeouts for standard benchmarks (up to 45 iterations)
  • Frontier models get appropriate extra time

🧪 Test Plan

describe('Analyzer Timeout Progression', () => {
  test('basic run uses minimum timeout', async () => {
    const runData = { iterations: 10, contextTokens: 1000, model: { size: 'standard' } };
    expect(calculateTimeout(runData)).toBe(120000);
  });

  test('complex run scales appropriately', async () => {
    const runData = { iterations: 45, contextTokens: 8000, model: { size: 'standard' } };
    const timeout = calculateTimeout(runData);
    expect(timeout).toBeGreaterThan(120000);
    expect(timeout).toBeLessThan(300000);
  });

  test('frontier models get bonus time', async () => {
    const standard = { iterations: 45, contextTokens: 8000, model: { size: 'standard' } };
    const frontier = { iterations: 45, contextTokens: 8000, model: { size: 'frontier' } };
    
    expect(calculateTimeout(frontier)).toBeGreaterThan(calculateTimeout(standard));
  });

  test('timeout capped at maximum', async () => {
    const extreme = { iterations: 50, contextTokens: 20000, model: { size: 'frontier' } };
    expect(calculateTimeout(extreme)).toBe(600000);
  });
});

📈 Expected Impact

  • Reduce timeout-related failures by 90%
  • Enable proper KSM scoring for all successful flag captures
  • Support analysis of complex attack patterns
  • Improve frontier model evaluation accuracy
  • Better user experience with progress indicators

🎯 Risk Assessment

  • Low risk: Longer timeouts don't affect result accuracy
  • Resource impact: Slightly increased LLM call costs (offset by fewer re-runs)
  • User impact: Better data, slight increase in analysis time

🔗 Related Issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions