Authors: Longyu Zhao (Staff Machine Learning Engineer), Gwendolyn Zhao (Staff Machine Learning Engineer), Peng Yan (Senior Machine Learning Engineer), Yuanlu Bai (Senior Machine Learning Engineer), Yuan Wang (Senior Machine Learning Engineer), Yao Cheng (Staff Machine Learning Engineer), Ang Xu (Principal Machine Learning Engineer), Zhaohong Han (Manager II, Ads Lightweight Ranking)
Introduction
Previously¹, we launched the next-generation serving stack for standard ads, which we call Nexus. Nexus decoupled candidate generation from scoring and moved us beyond the classic two-tower-only world, enabling richer model architectures while still meeting stringent latency and cost constraints.
Building on this system, we set out to design the first ads lightweight ranking model that goes beyond two towers. It jointly predicts three probabilities for each candidate ad: pCTR, the probability of a click; pGCTR30, the probability of a good click that lasts at least 30 seconds; and pOCTR, the probability of an outbound click to the advertiser’s destination. To support these objectives efficiently, we partition the query and Pin embeddings into task-specific CTR, gCTR30, and oCTR segments. For the CTR task, the fast two-tower prediction uses the first 64 dimensions of the CTR segment, while the three-tower prediction uses the full CTR segment together with richer cross features. For gCTR30 and oCTR tasks, full embeddings are shared between two-tower and three-tower predictions. This lets each task learn dedicated representations while sharing the overall model. In principle, Nexus places very few hard constraints on the architecture we can serve: cross-attention, sequence modeling, and more expressive interaction modules are all on the table.
However, in practice we quickly ran into the fundamental reality of ads lightweight ranking at Pinterest scale: for a typical request, we need to score on the order of hundreds of thousands of candidates (P99 post-targeting candidate counts can exceed 200K on some surfaces). We cannot simply keep increasing model complexity and expect to stay within our latency and cost budgets.
To strike a balance between latency and performance, we landed on a 3-tower co-train model design.
This design has a few key properties:
- We keep the query tower and Pin tower from the existing two-tower model, which lets us cache Pin embeddings offline and still obtain fast dot-product predictions for all candidates.
- We extend the architecture with a third cross tower that performs cross-attention between user sequences and candidate (Pin) features, plus an inter module that further mixes query, Pin, and cross embeddings.
- We co-train two-tower and three-tower predictions in a single model, giving us both fast but less accurate scores and slower but more accurate scores that we can deploy in different stages of the serving flow.
Put simply, the same model produces a fast two-tower score for every candidate and a richer three-tower score for a selected subset, so we can spend additional compute where it has the greatest impact.
Later in this blog, we will walk through the model architecture (cross tower and inter module), serving performance optimizations, and the two-stage scoring flow that leverages both two-tower and three-tower predictions.
By combining these changes, we maintained two-tower prediction quality while achieving around 30% reduction in offline loss for three-tower predictions across our engagement tasks, compared to the existing production model (details see below Offline Performance section). These offline gains translated into online lifts in CTR and gCTR30, reductions in cost per click (CPC), with a modest increase in infrastructure cost.
Model architecture
Our starting point was the existing two-tower engagement model, with separated query and Pin towers whose dot product feeds into task-specific heads. On top of this, we introduced two major components:
- A new cross tower that uses reduced-query cross-attention between user sequences and candidate features to capture high-order interactions.
- An inter module that jointly processes the query, Pin, and cross embeddings and produces a shared representation for all engagement tasks.
Below we describe the main design choices and trade-offs in each part.
Cross tower
The cross tower is responsible for modeling rich interactions between a user’s recent activity and a candidate ad. We use three on-site user sequence features (organic engagement, ads engagement, and search history) and four candidate features (advertiser ID, campaign ID, GraphSAGE embeddings, and Pin PinnerSAGE embeddings).
A natural first idea would be to build increasingly complex attention modules over these sequences and candidates. In practice, we explored several options:
- Merging full user sequences (across surfaces) and then running cross-attention with candidate features.
- Using shorter, truncated sequences to reduce compute and memory.
- Replacing attention with simpler interaction functions such as DIN-style pooling or average pooling.
- Crossing each sequence attribute with its corresponding candidate attribute individually, rather than merging first.
These variants exposed a clear trade-off: more expressive attention patterns (longer sequences, more attributes, per-attribute crossing) tended to improve offline loss but also increased latency, especially at high candidate counts. For example, using more complex cross architectures could reduce loss by several additional percentage points, but at the cost of tens of milliseconds of extra latency per request at 100K candidates.
We ultimately converged on an architecture that uses candidate side features to generate query tokens, which then interacts with user sequences to calculate attention. In this way, we can pick the most important candidate features and control the cost of transformer computation. This design gives us:
- Strong offline performance improvements versus production.
- A predictable compute profile that is easier to optimize and scale.
- A good balance between modeling capacity and serving latency.
The output of this cross tower is a cross embedding that summarizes how a user’s recent behavior interacts with a particular candidate ad.
Inter module
The inter module takes three inputs: the query embedding, the Pin embedding, and the cross embedding from the cross tower. Its goal is to produce a compact, shared representation that works well for all three engagement tasks (CTR, gCTR30, and oCTR), while keeping parameter count and serving latency under control.
Here as well, we evaluated multiple architectures:
- A deep MLP with DCN (Deep & Cross Network) layers.
- A standard MMoE (mixture-of-experts) with DCN.
- A top-K MMoE with DCN.
- A shared-bottom MLP with additive task-specific biases.
More complex structures such as MMoE with DCN achieved stronger loss reductions but also introduced noticeably higher latency compared to production. The shared-bottom MLP design provided a sweet spot: it delivered most of the performance gains while adding only modest latency, and it is architecturally simpler to optimize further.
In the final design, the inter module:
- Learns a shared logit for each task from the concatenated query, Pin, and cross embeddings.
- Adds task-specific biases computed from query and Pin embeddings for gCTR30 and oCTR.
- Outputs task logits that are then passed through sigmoid functions to produce probabilities.
This structure allows us to capture shared patterns across tasks while preserving enough task-specific flexibility.
Loss function
We train the model using a multi-task loss that combines main losses for the three-tower predictions with auxiliary losses for the two-tower co-train task.
The final loss takes the form of a weighted sum:
- Main three-tower losses for CTR, gCTR30, and oCTR.
- Co-train two-tower losses for CTR, gCTR30, and oCTR, computed from query–Pin dot products. For CTR, this fast auxiliary prediction uses the first 64 dimensions of the CTR embedding.
We tuned the task weights to balance learning stability and final performance, and landed on the following weighting scheme:
- Strong emphasis on main CTR loss.
- Moderate weight on main gCTR30 loss.
- Lower weight on main oCTR loss.
- Non-trivial but smaller weights on each of the co-train losses.
This configuration made the three-tower predictions the primary optimization target, while keeping the two-tower co-train task healthy enough to match or slightly improve on the production two-tower model.
The query tower produces a 192-dimensional task embedding: 144 dimensions for CTR, 32 for gCTR30, and 16 for oCTR. The Pin tower produces the corresponding task embedding and appends a 256-dimensional candidate-feature projection for the cross tower, producing a 448-dimensional Pin representation. For fast two-tower CTR scoring, we use only the first 64 dimensions of the 144-dimensional CTR segment. For three-tower CTR scoring, the inter module uses the full CTR segment together with the cross embedding. The additional Pin projection is computed in the Pin tower and cached offline, so the three-tower path can use these candidate features without per-candidate preprocessing at serving time.
Formally, our final loss is:
L = (100 * L_main_ctr + 5 * L_main_gctr30 + 1 * L_main_octr) + (10 * L_cotrain_ctr + 5 * L_cotrain_gctr30 + 1 * L_cotrain_octr)
Offline performance
We evaluated offline performance on held-out standard-ads data across three tasks (CTR, gCTR30, and oCTR), reporting relative loss reduction versus the production two-tower engagement model for both main (three-tower) and co-train (two-tower) predictions. The table reports ranges because we evaluated the model across multiple log sources; each endpoint is the result observed for a different source. For gCTR30, we observed a small degradation on the co-train loss which we deemed acceptable given the main task gains.
Across tasks, we observed:
Taken together, these results show that the co-train task maintains performance comparable to the existing production two-tower model, while the three-tower predictions deliver substantial improvements. This is important operationally: we can deprecate the standalone production two-tower engagement model and rely on the co-train head for fast scoring, without sacrificing quality.
Serving optimization
Serving a three-tower model over hundreds of thousands of candidates per request is expensive. To make the launch feasible, we invested heavily in model-level latency optimizations. Below are several techniques that had meaningful impact. Together, these changes reduced P99 model inference from over 200 ms to about 30 ms, leaving the final model only 1–2 ms slower than production.
Optimization 1: Move Pin pre-processing into the Pin tower
To perform cross-attention, sequence and candidate features must share the same dimensionality. In an initial design, we handled this with on-the-fly MLPs in the cross module, which added per-request compute proportional to the number of candidates.
Instead, we moved this preprocessing into the Pin tower. We append the processed Pin features to the original Pin embedding, increasing its dimension by an additional 256, and cache the resulting embedding offline. At serving time, the cross module can directly consume these enriched Pin embeddings with no additional per-request MLPs.
This change saved roughly 2 ms of latency at 100K candidates in our benchmarks.
Optimization 2: Lower precision
We also explored reduced-precision inference. By switching from FP32 to BF16 in the three-tower path, we significantly reduced model inference time while keeping model quality neutral.
On one representative benchmark, we observed:
- At 50K candidates, latency dropped from around 103 ms in FP32 to about 56 ms in BF16.
- At 20K candidates, latency dropped from around 45 ms to about 26 ms.
These gains played a key role in making the three-tower path practical at high candidate volumes.
Optimization 3: Late expansion of user features
In the three-tower engagement model, we compute predictions between one user and tens of thousands of candidate ads at once. User features are computed once and then expanded to match the batch size of candidates.
Earlier, this expansion happened just before the cross module to avoid duplicated computation. We realized we could delay expansion even further: instead of expanding before building the attention keys and values, we expand inside the cross module right before attention is computed.
This avoids redundant computation on large tensors and yields latency savings of around 15 ms at 100K candidates in our benchmarks.
Optimization 4: Pre-layer normalization in attention
Finally, we revisited how we apply LayerNorm inside the cross-attention module. Previously, we normalized the larger output sequence after attention, which has a batch size proportional to the number of candidates. We switched to normalizing the input sequence before attention instead; during serving this input has batch size 1, so the normalization work is much lower and independent of how many candidates we score.
During training, both options behave similarly. During serving, however, pre-layer normalization dramatically reduces the amount of work we do at large candidate counts. Flipping pre_lnorm from False to True reduced latency by about 3 ms at 100K candidates in our benchmarks.
Rethinking the serving flow
Even with model-level optimizations, running the full three-tower model on every candidate would still be too expensive. Post-targeting candidate counts can exceed 100K at P90 and reach up to over 200K at P99 on some surfaces. In early experiments where we scored all candidates with the three-tower model, model inference P99 latency exceeded 70 ms which is our timeout cutoff.
To tackle this, we redesigned the serving flow as a two-stage scoring pipeline that leverages both two-tower and three-tower predictions.
Stage 1: Fast scoring for all candidates
In the first stage, we use the two-tower head from the co-train model to score all candidates. These scores are combined into an initial utility: the overall ranking score that estimates a candidate ad’s value for the request and determines which candidates survive for later selection. This follows the existing production setup, but is powered by the new co-train model.
This stage is fast and inexpensive enough to run on the full candidate set.
Stage 2: Focused refinement with three-tower scoring
In the second stage, we identify the top-K candidates by utility and rescore only this subset with the three-tower model. Candidates outside this utility topK bypass the three-tower path and retain their two-tower scores.
We introduced a new hyperparameter, utility topK, which controls how many candidates the three-tower model sees. We tested several choices of the utility topK and measured both latency and downstream metrics such as clickthrough and web conversion impressions.
The trade-offs we observed:
- Smaller utility topK values reduce latency but can hurt web conversion impressions, because the final top-K selection stage needs a sufficiently large pool to satisfy different campaign groups and deduplication constraints.
- Very large utility topK values allow more candidates into the three-tower stage but directly increase latency, so we needed to pick a value that balanced candidate coverage and serving cost.
- Retaining candidates outside utility topK for later stages mitigates mixshifts with only a small additional latency cost.
We ultimately chose a utility topK of 40K. This threshold roughly corresponds to the 70th percentile for Home Feed and Related Pins and the 90th percentile for Search, which means the three-tower model scores the majority of candidates on most requests while keeping P99 latency within budget.
Final selection and bias considerations
After we obtain updated utility scores from the three-tower model for the utility topK subset, we blend them with the two-tower dot-product utilities. The final topK selection step then runs on this blended set of scores, selects pre-defined quotas from several candidate sources and performs deduplication to select the final set of ads shown to the user.
We made two design choices which may create concerns:
- Heuristic selection of the utility topK subset for three-tower scoring.
- Blending two-tower and three-tower utilities before final top-K selection.
Both heuristics are designed to favor high-utility candidates. In principle, they could bias the system toward items that already scored well in the two-tower stage, especially if three-tower predictions further amplify those scores.
To monitor this, we looked at calibration and mixshift. Encouragingly, we observed that the three-tower model actually reduces over-calibration for auction candidates, moving predicted CTR closer to realized CTR. We also did not observe a drop in standard web conversion impressions when using the blending option.
Online results and cost
In online A/B experiments on standard ads, the three-tower co-train model delivered around 1% gains in CTR, gCTR30, and oCTR, along with nearly 1% reductions in CPC, with only a modest increase in GPU spend.
Taken together, this represents a strong trade-off: meaningful engagement and efficiency gains for advertisers and users, with a small and well-understood increase in infra cost.
Conclusion
In this post, we walked through how we took Nexus beyond two towers by introducing a three-tower engagement co-train model for standard ads. On the modeling side, the cross tower and inter module allow us to capture richer interactions between user behavior and candidate ads. On the systems side, a combination of model-level optimizations and a two-stage scoring flow let us deploy this more powerful architecture while staying within tight latency and cost constraints.
Looking ahead, this work opens up several promising directions:
- Extending similar three-tower and co-train ideas to other objectives beyond engagement.
- Exploring even richer sequence modeling and attention patterns now that we have a scalable framework for late-stage scoring.
- Further tightening the feedback loop between offline architecture exploration, online performance, and infra-aware serving design.
Most importantly, it demonstrates that with the right system abstractions, we can continue to innovate on model architectures without losing sight of real-world constraints.
Acknowledgements
We thank Qingyu Zhou, Yuchen Shen, Li-Chien Lee, Qingmengting Wang, Zhixuan Shao, Tristan Nee, Sihan Wang, Lida Li, and Nuo Dou for their contributions to this project, and Renjun Zheng and Jamieson Kerns for their leadership support.
References
¹Yang, Xiao, et al. “Beyond Two Towers: Re-architecting the Serving Stack for Next-Gen Ads Lightweight Ranking Models (Part 1).” Pinterest Engineering Blog
Beyond Two Towers: Launching the 3-Tower Engagement Co-Train Model (Part 2) was originally published in Pinterest Engineering Blog on Medium, where people are continuing the conversation by highlighting and responding to this story.