Meta has detailed the infrastructure used to train GEM, its advertising recommendation foundation model. The system combines several thousand recent GPUs, trillions of sparse embedding parameters and billions of dense parameters. Its technical interest extends beyond ads: it shows how large-language-model methods change when users, objects and events dominate the data.
The short answer
| Question | Answer |
|---|---|
| Is GEM a chatbot? | No. It predicts and ranks advertising recommendations. |
| Why so many embeddings? | They represent a huge number of discrete identities and interactions. |
| What utilization does Meta report? | 20 to 25% MFU after doubling end-to-end training efficiency. |
| Are the benchmarks independent? | No. The figures and comparisons come from Meta. |
| Which ideas transfer elsewhere? | End-to-end profiling, lower precision, tailored communication and topology-aware placement. |
A hybrid model unlike a conventional LLM
An LLM mainly processes dense layers and token sequences. A recommendation system must also represent an enormous population of users, ads, categories and interactions. GEM therefore combines a dense component with billions of parameters and sparse tables reaching trillions.
That structure moves the bottleneck. Work is not only matrix computation; it includes irregular embedding access and heavy GPU communication. Adding accelerators without redesigning placement can increase waiting faster than throughput.
Attention and communication optimized together
Meta describes custom kernels including JFA, GDPA and BlockAttention for recommendation data. Its GDPA implementation is said to be twice as fast in forward processing and 1.6 times in backward processing over its prior baseline, reaching up to 3.5 times FA4 performance in some short-key/value cases.
Those comparisons belong to Meta's workloads and hardware. They still illustrate a useful rule: selecting a kernel because it leads a standard LLM benchmark may fail when sequence shape and sparsity differ.
Parallelism spans five dimensions with topology awareness. Dense layers, embeddings and communication should not cross the same links in the same way. Placement becomes part of the algorithm.
Low precision as a system strategy
GEM uses MXFP8 to reduce memory and communication while retaining required quality. Lower precision is more than a type switch: engineers must choose where to retain accuracy, how to scale blocks and how to detect numerical drift.
Meta says it doubled end-to-end efficiency to 20–25% MFU while quadrupling model FLOPs over twelve months. MFU measures use of theoretical accelerator capacity; a modest-looking figure can be meaningful on a communication- and sparse-access-heavy model.
Lessons for smaller infrastructure
Few teams will train at this scale, but the method transfers. Measure compute, data and network separately. Optimize the dominant path rather than the most visible kernel. Validate low precision on a business metric and place data for the real topology.
Cost is another lesson. Quadrupling FLOPs matters only when recommendation gains justify energy, hardware and complexity. Infrastructure metrics should remain tied to a product result.
GEM shows recommendation models absorbing LLM techniques without becoming LLMs. Their distinct challenge remains combining dense compute, gigantic sparse memory and a network able to feed both.




Join the discussion
Comments
Loading comments…