Skip to content

Add real-time cost tracking dashboard and ROI analysis #77

Description

@r3y3r53

title: Add real-time cost tracking dashboard and ROI analysis
labels: enhancement, ux, monitoring, cost

🚨 Problem

Complete lack of cost visibility during benchmark runs:

  • Can't monitor budget in real-time ("blind money spending")
  • No cost-per-run estimates before starting batches
  • No cumulative spend tracking across benchmarks
  • No cost efficiency comparisons between model families
  • Risk of unexpected charges with large batch runs

📊 Current State

Users run blind:

$ oasis run --batch 7x14 --models all
[Progress] Model 3/7, Lab 8/14... 
# No cost awareness! Is this $5 or $50 or $500?

$ oasis report
📊 Results: GLM-5.3: 10/14, V4.1-Flash: 9/14...
# But no cost data included!

🔍 Evidence

Our 175+ run benchmark cost:

  • GLM-5.3: ~$0.014/run (cheap flash class)
  • Qwen3.8-27B: ~$0.089/run (larger model, same flag rate)
  • Total unknown until we manually calculated from logs

Users asking for "how much will this cost?" get silence.

🛠️ Proposed Solution

1. Cost Calculation Engine

// src/cost/cost-calculator.ts
export interface CostConfig {
  [modelId: string]: {
    input: number;    // per 1M tokens
    output: number;   // per 1M tokens
    currency: string;
  };
}

export const DEEPINFRA_COSTS: CostConfig = {
  'zai-org/GLM-5.3': {
    input: 0.14,  // $0.14/M input tokens
    output: 0.42, // $0.42/M output tokens  
    currency: 'USD'
  },
  'thinkingmachines/Inkling-Small': {
    input: 0.18,
    output: 0.54,
    currency: 'USD'
  },
  'deepseek-ai/DeepSeek-V4.1-Flash': {
    input: 0.19,
    output: 0.57,
    currency: 'USD'
  },
  'Qwen/Qwen3.8-27B': {
    input: 0.27,  // Larger model, higher cost
    output: 0.81,
    currency: 'USD'
  }
  // ... all model costs
};

export class CostCalculator {
  calculateRunCost(model: string, inputTokens: number, outputTokens: number): number {
    const config = this.costs[model];
    if (!config) return 0; // Unknown model
    
    const inputCost = (inputTokens / 1_000_000) * config.input;
    const outputCost = (outputTokens / 1_000_000) * config.output;
    return inputCost + outputCost;
  }

  estimateRunCost(model: string, maxIterations: number): number {
    // Historical averages from our data
    const avgInputTokens = 45000;  // per 45 iterations
    const avgOutputTokens = 2500;  // per run
    
    return this.calculateRunCost(model, avgInputTokens, avgOutputTokens);
  }

  estimateBatchCost(modelRuns: ModelRun[]): BatchCostEstimate {
    const estimated = modelRuns.map(run => ({
      model: run.model,
      estimatedCost: this.estimateRunCost(run.model, run.maxIterations),
      runs: run.count
    }));

    return {
      totalCost: estimated.reduce((sum, e) => sum + (e.estimatedCost * e.runs), 0),
      byModel: estimated,
      currency: 'USD'
    };
  }
}

2. Real-Time Cost Dashboard

// src/ui/cost-dashboard.ts
export class CostDashboard {
  renderLiveProgress(currentRun: CurrentRun, batchStats: BatchStats): string {
    const spent = batchStats.totalSpent;
    const budget = batchStats.budget || Infinity;
    const remaining = Math.max(0, budget - spent);
    const progress = budget !== Infinity ? (spent / budget) * 100 : 0;

    return this.formatProgress({
      current: {
        model: currentRun.model,
        iterations: `${currentRun.iterations}/${currentRun.maxIterations}`,
        estCost: `$${currentRun.estimatedCost.toFixed(4)}`,
        currentCost: `$${currentRun.cost.toFixed(4)}`
      },
      totals: {
        spent: `$${spent.toFixed(2)}`,
        remaining: `$${remaining.toFixed(2)}`,
        budget: budget === Infinity ? '∞' : `$${budget.toFixed(2)}`,
        progress: `${progress.toFixed(1)}%`
      },
      batch: {
        current: `${batchStats.currentRun}/${batchStats.totalRuns}`,
        eta: this.format_eta(batchStats.eta),
        avgCost: `$${batchStats.avgCost.toFixed(3)}/run`,
        total: `$${batchStats.totalEstimated.toFixed(2)}`
      }
    });
  }

  formatProgress(data: ProgressData): string {
    return `
╔════════════════════════════════════════════════════════════╗
║                    OASIS COST TRACKER                        ║
╠════════════════════════════════════════════════════════════╣
║ Current Run: ${data.current.model.padEnd(20)} (${data.current.iterations.padEnd(10)})      ║
║ Est. Cost: ${data.current.estCost.padEnd(10)} | Current: ${data.current.currentCost.padEnd(10)}            ║
════════════════════════════════════════════════════════════
║ TOTAL SPENT:    ${data.totals.spent.padEnd(8)} (${data.totals.progress.padEnd(4)})                ║
║ BUDGET:         ${data.totals.budget.padEnd(8)}                              ║
║ REMAINING:      ${data.totals.remaining.padEnd(8)}                          ║
════════════════════════════════════════════════════════════
║ Run ${data.batch.current.padEnd(4)} │ ETA: ${data.batch.eta.padEnd(8)}  │ Budget: ${(parseFloat(data.totals.remaining) / parseFloat(data.totals.spent) * 100).toFixed(0)}% left     ║
║ Cost/run: ${data.batch.avgCost.padEnd(8)} │ Total: ${data.batch.total.padEnd(10)}         ║
╚══════════════════════════════════════════════════════════╝
    `.trim();
  }
}

3. Cost-Aware CLI Integration

// src/commands/run.ts
export async function runCommand(options: RunOptions): Promise<void> {
  // Pre-cost estimation for batch runs
  if (options.batch) {
    const modelRuns = parseBatchConfig(options);
    const costEstimate = calculator.estimateBatchCost(modelRuns);
    
    console.log(`Estimated batch cost: $${costEstimate.totalCost.toFixed(2)}`);
    console.log(`
Cost Breakdown:
${costEstimate.byModel.map(m => `  ${m.model.padEnd(30)} $${(m.estimatedCost * m.runs).toFixed(2)} (${m.runs} runs)`).join('\n')}
    `);

    if (options.budget && costEstimate.totalCost > options.budget) {
      throw new Error(`Estimated cost $${costEstimate.totalCost.toFixed(2)} exceeds budget $${options.budget.toFixed(2)}`);
    }

    // Confirm unless --no-confirm
    if (!options.noConfirm) {
      const proceed = await confirm('Continue with the benchmark?');
      if (!proceed) process.exit(0);
    }
  }

  // Start with cost tracking
  const costTracker = new CostTracker();
  const batchStart = Date.now();
  
  for (const run of batchRuns) {
    const runId = await costTracker.startRun(run.model, run.maxIterations);
    
    // ... existing run logic with cost updates
    await runWithCostTracking(run, (progress) => {
      costTracker.updateUsage(runId, progress.inputTokens, progress.outputTokens);
      updateDashboard({
        current: progress,
        stats: costTracker.getBatchStats(),
        budget: options.budget
      });
    });
  }
  
  // Final cost report
  const finalCosts = await costTracker.finalizeBatch();
  console.log(new CostReport(finalCosts).render());
}

4. Cost-Aware Reporting

// src/reports/cost-report.ts
export interface CostReport {
  totalCost: number;
  modelCosts: ModelCostMap;
  costPerFlag: ModelROI[];
  roiRanking: ModelROI[];
  recommendations: string[];
}

export class CostReporter {
  generate(benchmarkResults: BenchmarkResults, costData: CostData): CostReport {
    return {
      totalCost: costData.totalSpent,
      modelCosts: this.calculateModelCosts(costData),
      costPerFlag: this.calculateCostPerFlag(benchmarkResults, costData),
      roiRanking: this.rankByROI(benchmarkResults, costData),
      recommendations: this.generateRecommendations(benchmarkResults, costData)
    };
  }

  calculateCostPerFlag(results: BenchmarkResults, costs: CostData): ModelROI[] {
    return Object.entries(results)
      .map(([model, result]: [string, ModelResult]) => {
        const totalCost = costs.byModel[model] || 0;
        const costPerFlag = result.flags > 0 ? totalCost / result.flags : Infinity;
        
        return {
          model,
          flags: result.flags,
          totalCost,
          costPerFlag,
          avgKSM: result.avgKsm
        };
      })
      .sort((a, b) => a.costPerFlag - b.costPerFlag);
  }

  render(report: CostReport): string {
    return `
═══════════════════════════════════════════════════════════════
                    COST ANALYSIS REPORT
═══════════════════════════════════════════════════════════════
TOTAL SPENT: $${report.totalCost.toFixed(2)}

TOP ROI MODELS (Cost per Flag Captured):
${report.costPerFlag.slice(0, 5)
  .map(m => `• ${m.model.padEnd(20)} $${m.costPerFlag.toFixed(3)}/flag (${m.flags} flags)`)
  .join('\n')}

COST BREAKDOWN BY MODEL:
${Object.entries(report.modelCosts)
  .sort(([,a], [,b]) => b - a)
  .map(([model, cost]) => `• ${model.padEnd(20)} $${cost.toFixed(2)}`)
  .join('\n')}

💡 RECOMMENDATIONS:
${report.recommendations.map(r => `• ${r}`).join('\n')}
    `.trim();
  }
}

✅ Acceptance Criteria

  • Real-time cost display during runs (progress bars with cost)
  • Cost estimation before starting benchmarks
  • Budget tracking with alerts (optional budget limits)
  • Per-model cost efficiency analysis (cost/flag)
  • Cost data included in final benchmark reports
  • CSV export of full cost breakdown
  • CLI warnings for unexpectedly expensive runs

🧪 Test Plan

describe('Cost Tracking', () => {
  test('accurate cost calculation', () => {
    const cost = calculator.calculateRunCost('zai-org/GLM-5.3', 50000, 3000);
    expect(cost).toBeCloseTo(8.76, 2); // 50K*0.14 + 3K*0.42 = $7.00 + $1.26
  });

  test('real-time updates', async () => {
    const tracker = new CostTracker();
    const runId = await tracker.startRun('zai-org/GLM-5.3', 45);
    
    await tracker.updateUsage(runId, 10000, 500);
    const stats = tracker.getBatchStats();
    
    expect(stats.totalSpent).toBeCloseTo(1.70, 2); // 10K*0.14 + 0.5K*0.42
  });

  test('budget alerts', async () => {
    const tracker = new CostTracker({ budget: 10 });
    await tracker.updateBatchCost(9.50);
    await tracker.updateBatchCost(0.51); // Should trigger alert
    
    expect(tracker.getBatchStats().budgetAlert).toBe(true);
  });
});

📋 Configuration

# config/cost.yml
costs:
  tracking:
    enabled: true
    real_time: true
    budget_alerts: true
  currency: USD
  models:
    zai-org/GLM-5.3: { input: 0.14, output: 0.42 }
    thinkingmachines/Inkling-Small: { input: 0.18, output: 0.54 }
    deepseek-ai/DeepSeek-V4.1-Flash: { input: 0.19, output: 0.57 }
    Qwen/Qwen3.8-27B: { input: 0.27, output: 0.81 }
    # ... update with current DeepInfra pricing
  
  dashboard:
    show_budget: true
    show_eta: true
    update_interval: 5s

📊 Sample Output

OASIS Batch Run - Cost Tracker v0.1
══════════════════════════════════════════════════════════════
Current Run: V4.1-Flash (12/45) | Est. Cost: $0.0043 | Current: $0.0037
═════════════════════════════════════════════════════════════
TOTAL SPENT:    $12.34 (37% of budget) | REMAINING: $20.66
	Run 12/98 │ ETA: 45m | Budget: 63% left
Cost/run: $0.016 │ Total: $12.34

📈 Expected Impact

  • Enable budget-conscious benchmarking decisions
  • Objective model efficiency comparisons (cost !== quality)
  • Prevent unexpected charges on large batches
  • Make cost part of model selection criteria
  • Professional-grade cost reporting for enterprise users

🎯 Use Cases Enabled

  • "How much will this 7x14 batch cost?" ✅
  • "Which model is most cost-effective?" ✅
  • "Stop if we exceed $X budget" ✅
  • "Export cost data to finance" ✅

🔗 Related Issues

  • Addresses user requests for cost visibility
  • Enables enterprise adoption with budget controls
  • Complements benchmark quality metrics with cost efficiency

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions