Clipping versus compressing
The lazy approach clips: any colour outside the target gamut is moved to the nearest point on its boundary. That works until two distinguishable colours land on the same boundary point, at which point the detail between them disappears. A saturated red flower photographed against saturated red fabric becomes one flat shape.
Compressing instead shifts the whole range inward, giving up a little accuracy in colours that were already reproducible to keep the out-of-gamut ones distinct. Which is correct depends on the content, which is why the choice is a grading decision rather than a setting.
Why it comes up in HDR work
BT.2020 is much wider than BT.709, and no consumer display covers it fully. Every HDR playback chain is therefore doing gamut mapping, usually invisibly, in the display's own processing.
Going the other direction, converting an HDR master down to an SDR deliverable, you are doing it explicitly, and doing it badly is the other half of why converted files look wrong. Tone mapping handles the brightness; gamut mapping handles the colour; skipping either produces a bad result for different reasons.
Questions people ask
Do I need to do gamut mapping myself?
For playback, no: displays do it continuously and invisibly. You do it deliberately only when producing a deliverable for a narrower colour space, which in practice means an SDR export from an HDR master. Converting HDR to SDR covers that path.
Related: converting HDR to SDR.