Skip to content

Design: evaluation loop — performance data as reward signal #5

Description

@gidutz

Adtech has unusually clean ground truth: ROAS, CTR, fill rate, viewability are measurable within hours. We should exploit this.

Components:

  • Performance data ingest (from ad platforms back into our store)
  • Reward signal definition per agent role (what does "good" mean?)
  • Offline eval harness (replay historical decisions)
  • Online eval / shadow mode (run agent in parallel with a human, compare outcomes)
  • Regression tracking across agent versions

Open questions:

  • Attribution lag — how long do we wait before scoring a decision?
  • Confounders (seasonality, exogenous events, creative fatigue)
  • Per-account vs. global eval

Deliverable: eval harness spec plus initial metric definitions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    pillar:evalsPerformance feedback loop and metricstype:designDesign discussion

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions