title: Add real-time cost tracking dashboard and ROI analysis
labels: enhancement, ux, monitoring, cost
🚨 Problem
Complete lack of cost visibility during benchmark runs:
- Can't monitor budget in real-time ("blind money spending")
- No cost-per-run estimates before starting batches
- No cumulative spend tracking across benchmarks
- No cost efficiency comparisons between model families
- Risk of unexpected charges with large batch runs
📊 Current State
Users run blind:
$ oasis run --batch 7x14 --models all
[Progress] Model 3/7, Lab 8/14...
# No cost awareness! Is this $5 or $50 or $500?
$ oasis report
📊 Results: GLM-5.3: 10/14, V4.1-Flash: 9/14...
# But no cost data included!
🔍 Evidence
Our 175+ run benchmark cost:
- GLM-5.3: ~$0.014/run (cheap flash class)
- Qwen3.8-27B: ~$0.089/run (larger model, same flag rate)
- Total unknown until we manually calculated from logs
Users asking for "how much will this cost?" get silence.
🛠️ Proposed Solution
1. Cost Calculation Engine
// src/cost/cost-calculator.ts
export interface CostConfig {
[modelId: string]: {
input: number; // per 1M tokens
output: number; // per 1M tokens
currency: string;
};
}
export const DEEPINFRA_COSTS: CostConfig = {
'zai-org/GLM-5.3': {
input: 0.14, // $0.14/M input tokens
output: 0.42, // $0.42/M output tokens
currency: 'USD'
},
'thinkingmachines/Inkling-Small': {
input: 0.18,
output: 0.54,
currency: 'USD'
},
'deepseek-ai/DeepSeek-V4.1-Flash': {
input: 0.19,
output: 0.57,
currency: 'USD'
},
'Qwen/Qwen3.8-27B': {
input: 0.27, // Larger model, higher cost
output: 0.81,
currency: 'USD'
}
// ... all model costs
};
export class CostCalculator {
calculateRunCost(model: string, inputTokens: number, outputTokens: number): number {
const config = this.costs[model];
if (!config) return 0; // Unknown model
const inputCost = (inputTokens / 1_000_000) * config.input;
const outputCost = (outputTokens / 1_000_000) * config.output;
return inputCost + outputCost;
}
estimateRunCost(model: string, maxIterations: number): number {
// Historical averages from our data
const avgInputTokens = 45000; // per 45 iterations
const avgOutputTokens = 2500; // per run
return this.calculateRunCost(model, avgInputTokens, avgOutputTokens);
}
estimateBatchCost(modelRuns: ModelRun[]): BatchCostEstimate {
const estimated = modelRuns.map(run => ({
model: run.model,
estimatedCost: this.estimateRunCost(run.model, run.maxIterations),
runs: run.count
}));
return {
totalCost: estimated.reduce((sum, e) => sum + (e.estimatedCost * e.runs), 0),
byModel: estimated,
currency: 'USD'
};
}
}
2. Real-Time Cost Dashboard
// src/ui/cost-dashboard.ts
export class CostDashboard {
renderLiveProgress(currentRun: CurrentRun, batchStats: BatchStats): string {
const spent = batchStats.totalSpent;
const budget = batchStats.budget || Infinity;
const remaining = Math.max(0, budget - spent);
const progress = budget !== Infinity ? (spent / budget) * 100 : 0;
return this.formatProgress({
current: {
model: currentRun.model,
iterations: `${currentRun.iterations}/${currentRun.maxIterations}`,
estCost: `$${currentRun.estimatedCost.toFixed(4)}`,
currentCost: `$${currentRun.cost.toFixed(4)}`
},
totals: {
spent: `$${spent.toFixed(2)}`,
remaining: `$${remaining.toFixed(2)}`,
budget: budget === Infinity ? '∞' : `$${budget.toFixed(2)}`,
progress: `${progress.toFixed(1)}%`
},
batch: {
current: `${batchStats.currentRun}/${batchStats.totalRuns}`,
eta: this.format_eta(batchStats.eta),
avgCost: `$${batchStats.avgCost.toFixed(3)}/run`,
total: `$${batchStats.totalEstimated.toFixed(2)}`
}
});
}
formatProgress(data: ProgressData): string {
return `
╔════════════════════════════════════════════════════════════╗
║ OASIS COST TRACKER ║
╠════════════════════════════════════════════════════════════╣
║ Current Run: ${data.current.model.padEnd(20)} (${data.current.iterations.padEnd(10)}) ║
║ Est. Cost: ${data.current.estCost.padEnd(10)} | Current: ${data.current.currentCost.padEnd(10)} ║
════════════════════════════════════════════════════════════
║ TOTAL SPENT: ${data.totals.spent.padEnd(8)} (${data.totals.progress.padEnd(4)}) ║
║ BUDGET: ${data.totals.budget.padEnd(8)} ║
║ REMAINING: ${data.totals.remaining.padEnd(8)} ║
════════════════════════════════════════════════════════════
║ Run ${data.batch.current.padEnd(4)} │ ETA: ${data.batch.eta.padEnd(8)} │ Budget: ${(parseFloat(data.totals.remaining) / parseFloat(data.totals.spent) * 100).toFixed(0)}% left ║
║ Cost/run: ${data.batch.avgCost.padEnd(8)} │ Total: ${data.batch.total.padEnd(10)} ║
╚══════════════════════════════════════════════════════════╝
`.trim();
}
}
3. Cost-Aware CLI Integration
// src/commands/run.ts
export async function runCommand(options: RunOptions): Promise<void> {
// Pre-cost estimation for batch runs
if (options.batch) {
const modelRuns = parseBatchConfig(options);
const costEstimate = calculator.estimateBatchCost(modelRuns);
console.log(`Estimated batch cost: $${costEstimate.totalCost.toFixed(2)}`);
console.log(`
Cost Breakdown:
${costEstimate.byModel.map(m => ` ${m.model.padEnd(30)} $${(m.estimatedCost * m.runs).toFixed(2)} (${m.runs} runs)`).join('\n')}
`);
if (options.budget && costEstimate.totalCost > options.budget) {
throw new Error(`Estimated cost $${costEstimate.totalCost.toFixed(2)} exceeds budget $${options.budget.toFixed(2)}`);
}
// Confirm unless --no-confirm
if (!options.noConfirm) {
const proceed = await confirm('Continue with the benchmark?');
if (!proceed) process.exit(0);
}
}
// Start with cost tracking
const costTracker = new CostTracker();
const batchStart = Date.now();
for (const run of batchRuns) {
const runId = await costTracker.startRun(run.model, run.maxIterations);
// ... existing run logic with cost updates
await runWithCostTracking(run, (progress) => {
costTracker.updateUsage(runId, progress.inputTokens, progress.outputTokens);
updateDashboard({
current: progress,
stats: costTracker.getBatchStats(),
budget: options.budget
});
});
}
// Final cost report
const finalCosts = await costTracker.finalizeBatch();
console.log(new CostReport(finalCosts).render());
}
4. Cost-Aware Reporting
// src/reports/cost-report.ts
export interface CostReport {
totalCost: number;
modelCosts: ModelCostMap;
costPerFlag: ModelROI[];
roiRanking: ModelROI[];
recommendations: string[];
}
export class CostReporter {
generate(benchmarkResults: BenchmarkResults, costData: CostData): CostReport {
return {
totalCost: costData.totalSpent,
modelCosts: this.calculateModelCosts(costData),
costPerFlag: this.calculateCostPerFlag(benchmarkResults, costData),
roiRanking: this.rankByROI(benchmarkResults, costData),
recommendations: this.generateRecommendations(benchmarkResults, costData)
};
}
calculateCostPerFlag(results: BenchmarkResults, costs: CostData): ModelROI[] {
return Object.entries(results)
.map(([model, result]: [string, ModelResult]) => {
const totalCost = costs.byModel[model] || 0;
const costPerFlag = result.flags > 0 ? totalCost / result.flags : Infinity;
return {
model,
flags: result.flags,
totalCost,
costPerFlag,
avgKSM: result.avgKsm
};
})
.sort((a, b) => a.costPerFlag - b.costPerFlag);
}
render(report: CostReport): string {
return `
═══════════════════════════════════════════════════════════════
COST ANALYSIS REPORT
═══════════════════════════════════════════════════════════════
TOTAL SPENT: $${report.totalCost.toFixed(2)}
TOP ROI MODELS (Cost per Flag Captured):
${report.costPerFlag.slice(0, 5)
.map(m => `• ${m.model.padEnd(20)} $${m.costPerFlag.toFixed(3)}/flag (${m.flags} flags)`)
.join('\n')}
COST BREAKDOWN BY MODEL:
${Object.entries(report.modelCosts)
.sort(([,a], [,b]) => b - a)
.map(([model, cost]) => `• ${model.padEnd(20)} $${cost.toFixed(2)}`)
.join('\n')}
💡 RECOMMENDATIONS:
${report.recommendations.map(r => `• ${r}`).join('\n')}
`.trim();
}
}
✅ Acceptance Criteria
🧪 Test Plan
describe('Cost Tracking', () => {
test('accurate cost calculation', () => {
const cost = calculator.calculateRunCost('zai-org/GLM-5.3', 50000, 3000);
expect(cost).toBeCloseTo(8.76, 2); // 50K*0.14 + 3K*0.42 = $7.00 + $1.26
});
test('real-time updates', async () => {
const tracker = new CostTracker();
const runId = await tracker.startRun('zai-org/GLM-5.3', 45);
await tracker.updateUsage(runId, 10000, 500);
const stats = tracker.getBatchStats();
expect(stats.totalSpent).toBeCloseTo(1.70, 2); // 10K*0.14 + 0.5K*0.42
});
test('budget alerts', async () => {
const tracker = new CostTracker({ budget: 10 });
await tracker.updateBatchCost(9.50);
await tracker.updateBatchCost(0.51); // Should trigger alert
expect(tracker.getBatchStats().budgetAlert).toBe(true);
});
});
📋 Configuration
# config/cost.yml
costs:
tracking:
enabled: true
real_time: true
budget_alerts: true
currency: USD
models:
zai-org/GLM-5.3: { input: 0.14, output: 0.42 }
thinkingmachines/Inkling-Small: { input: 0.18, output: 0.54 }
deepseek-ai/DeepSeek-V4.1-Flash: { input: 0.19, output: 0.57 }
Qwen/Qwen3.8-27B: { input: 0.27, output: 0.81 }
# ... update with current DeepInfra pricing
dashboard:
show_budget: true
show_eta: true
update_interval: 5s
📊 Sample Output
OASIS Batch Run - Cost Tracker v0.1
══════════════════════════════════════════════════════════════
Current Run: V4.1-Flash (12/45) | Est. Cost: $0.0043 | Current: $0.0037
═════════════════════════════════════════════════════════════
TOTAL SPENT: $12.34 (37% of budget) | REMAINING: $20.66
Run 12/98 │ ETA: 45m | Budget: 63% left
Cost/run: $0.016 │ Total: $12.34
📈 Expected Impact
- Enable budget-conscious benchmarking decisions
- Objective model efficiency comparisons (cost !== quality)
- Prevent unexpected charges on large batches
- Make cost part of model selection criteria
- Professional-grade cost reporting for enterprise users
🎯 Use Cases Enabled
- "How much will this 7x14 batch cost?" ✅
- "Which model is most cost-effective?" ✅
- "Stop if we exceed $X budget" ✅
- "Export cost data to finance" ✅
🔗 Related Issues
- Addresses user requests for cost visibility
- Enables enterprise adoption with budget controls
- Complements benchmark quality metrics with cost efficiency
title: Add real-time cost tracking dashboard and ROI analysis
labels: enhancement, ux, monitoring, cost
🚨 Problem
Complete lack of cost visibility during benchmark runs:
📊 Current State
Users run blind:
🔍 Evidence
Our 175+ run benchmark cost:
Users asking for "how much will this cost?" get silence.
🛠️ Proposed Solution
1. Cost Calculation Engine
2. Real-Time Cost Dashboard
3. Cost-Aware CLI Integration
4. Cost-Aware Reporting
✅ Acceptance Criteria
🧪 Test Plan
📋 Configuration
📊 Sample Output
📈 Expected Impact
🎯 Use Cases Enabled
🔗 Related Issues