Pontiac Wiki Pontiac Wiki

Model Performance & Validation

Analytics Documentation

  • Overview
  • Model Scoring Overview
  • User Defined Models
    • Overview
    • New UDM Set Up
  • Model Performance & Validation
  • Planning
    • Planning Overview
    • New Plan Set Up
    • New Plan Draft
    • Plan Results
    • Plan to Action
  • Audience
    • Audience Overview
    • New Report Set Up
    • Report Results
  • Targeting
    • Targeting Overview
    • New Report Set Up
    • Report Results
  • Incrementality
    • Incrementality Overview
    • New Report Set Up
    • Report Results
  • Public Models
  • Analytics Model Association
  1. Home
  2. Analytics Documentation
  3. Model Performance & Validation

Model Performance & Validation

How well it has worked and how you know.

What the Model Does

The Pontiac targeting model scores bid requests using patterns learned from a campaign’s historical delivery and identifies inventory that shows stronger or weaker alignment with the campaign’s optimization objective, such as conversions or clicks.

The model does not directly change the line’s bid price. Instead, it changes which bid opportunities are prioritized and how often the line bids on them, concentrating more delivery toward higher-scoring inventory.

The model score is a ranking signal, not a predicted probability that an impression will convert or receive a click. For details on how model scores are created, matched, combined, and applied during bidding, see Model Scoring Overview.

Measured Performance

Historical validation shows that higher-scoring inventory produced stronger outcome rates than the campaign average.

Share of highest-scoring inventoryOutcome rate vs. campaign average
Top 10%2.33×
Top 25%1.96×
Top 50%1.53×

In other words, the model was able to rank inventory so that the strongest-scoring portion contained a disproportionate share of campaign outcomes.

The direction of the result was positive for every campaign included in the validation, across both conversion-optimized and click-optimized campaigns.

These results describe historical model performance. They do not guarantee that every future campaign or line will achieve the same lift.

How Performance Was Measured

The validation methodology is designed to test whether the model can rank future inventory rather than simply reproduce patterns from the data on which it was trained.

Trained on the Past, Tested on the Future

The model is built from an earlier period of campaign delivery and evaluated against a later period it has not seen during training.

This temporal validation tests whether patterns learned from earlier delivery continue to rank inventory effectively in subsequent delivery. It avoids evaluating the model against the same observations used to build it.

Compared Within the Same Day

Conversions may occur days after an impression is served. As a result, impressions delivered near the end of a measurement period can appear to perform worse simply because associated conversions have not yet occurred.

To reduce this distortion, performance comparisons are made between impressions served on the same day. Higher- and lower-scoring inventory therefore receive comparable conversion lookback time.

Evaluated Per Campaign

Campaigns are evaluated independently rather than pooling all impressions and outcomes into one aggregate result.

This prevents a single high-volume or high-performing campaign from dominating the overall measurement and makes it possible to evaluate whether the model ranks inventory consistently across campaigns.

The figures reported above are based on historical replay of real campaign delivery. Measurement of live campaign performance is ongoing.

How Model Concentration Is Controlled

The model does not have a fixed level of concentration. How aggressively it affects bidding is controlled through the model settings applied to each line.

In standard probability-based, or prob, output mode, the line uses a minimum threshold and a maximum threshold:

  • Scores at or below the minimum do not bid.
  • Scores at or above the maximum always bid.
  • Scores between the two thresholds bid with probability equal to the bidder score.

For example, a request scoring 0.70 within the configured threshold band bids approximately 70% of the time, while a request scoring 0.40 bids approximately 40% of the time.

Changing the thresholds therefore changes how strongly bidding is concentrated toward higher-scoring inventory without changing the underlying model scores themselves.

Raising the minimum suppresses more lower-scoring inventory. Lowering the maximum causes more higher-scoring inventory to always bid. Setting the minimum and maximum to the same value removes the probabilistic band and produces hard-cutoff behavior.

For the complete threshold logic, including neutral 0.5 behavior and classification (cls) mode, see Model Scoring Overview.

Concentration in Validation

The validation analysis also evaluated different levels of model concentration:

ScenarioBids placedPreferred inventoryDeprioritized inventory
Light47% of opportunities+11% vs. even spend−11%
Moderate43%+26%−26%
Aggressive41%+39%−39%

These labels describe the concentration scenarios evaluated during validation. They are not separate scoring models or output modes.

Increasing concentration shifts more bidding toward preferred inventory and away from inventory the model scores less favorably. In practice, lines can begin with a broader bidding strategy while delivery establishes and then be adjusted based on performance and available scale.

Where the Model Is Strongest

  • Campaigns With Meaningful Delivery History

    Pontiac Analytics models learn from observed campaign performance. More observed outcomes provide stronger evidence for distinguishing useful patterns from noise.

    Analytics reports therefore include a confidence rating based partly on the amount of supporting data. When the available evidence is insufficient, the report does not silently present a weak model as reliable.

    • Inventory With Strong Model Coverage

    The model can only influence a request when at least one model entry matches it.

    A request that matches no model entry receives the neutral bidder score of exactly 0.5, meaning the model has no positive or negative signal for that opportunity.

    Performance therefore depends partly on coverage: how often the attributes represented in the model are also available in incoming bid requests.

    Display and app inventory can provide strong coverage because attributes such as domain, app, publisher, exchange, device, and time are commonly available for scoring.

    Campaigns Where Inventory Performance Varies

    The model is designed to distinguish between stronger and weaker opportunities.

    When performance varies meaningfully across publishers, devices, geographies, times, or other supported attributes, the model has greater opportunity to concentrate bidding toward stronger inventory.

    If available inventory performs nearly identically, there is less meaningful variation for the model to exploit.

    Model Confidence

    Each Pontiac Analytics report includes a confidence rating of high, medium, low, or none.

    Confidence reflects three areas:

    • The number of observed conversions or clicks relative to measured minimums
    • The model’s cross-validated ability to rank outcomes
    • The amount of training data supporting each field

    The overall rating reflects the weakest of these components because insufficient support in any one area can undermine the reliability of the resulting model.

    For practical use:

    • Low — treat the model as directional. Do not scale spend based on the model alone.
    • None — no bid model was produced.

    User Defined Models do not receive a confidence rating because confidence represents a measurement of the data supporting an Analytics-generated model rather than a property of a hand-written score file.

    See Model Scoring Overview for additional detail on how Analytics-generated scores and field-importance weights are constructed.

    What the Results Do Not Claim

    The validation results should be interpreted within the scope of what was measured.

    Historical validation results do not guarantee the same performance for every campaign. Results depend on available training data, model confidence, request coverage, inventory variation, and campaign delivery conditions.

    These results measure attributed outcomes, using the same attribution basis as campaign reporting. They do not measure incremental lift. Establishing incrementality requires a controlled experiment.

    Historical replay measures whether the model successfully ranked inventory in held-out campaign data. It is not the same as a randomized live-campaign experiment.

    No comparison with the targeting or optimization models of other platforms is claimed or implied.

    A higher bidder score does not represent a predicted conversion or click probability. The model score ranks inventory according to the strength of the model’s signals.

    © 2026 Pontiac Wiki Sec · All Rights Reserved · Developed by RDK

    • Contact Us
    • Privacy Policy
    • Terms & Conditions
    • Pontiac.media