A study evaluated three commercial multimodal large language models (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) on 427 invasive breast cancer cases using histological images to determine if they could achieve the accuracy needed for treatment-determining diagnosis. Claude achieved the highest concordance in histologic-type classification but the lowest kappa value, misclassifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. None of the models reliably identified HER2-positive disease. Each vendor failed in a different direction: Claude and GPT-5.5 under-detected while Gemini over-called. Even 12 prompt variants (4,056 calls) did not improve sensitivity. The authors concluded that none of these models are ready for autonomous clinical deployment and their value is only assistive at a cost of USD 0.20-0.50 per case.