Method for this part
Two models run here. One returns 478 landmarks and says where the face is; the other labels every pixel as background, hair, body skin, face skin, clothes or accessory. The first decides the crop, the second decides which part each pixel belongs to.