The study evaluated the performance of five cloud-based large language models (Claude, DeepSeek, GPT, Grok, and Qwen) in extracting glaucoma information from clinical notes. The research included 1,250 patients from the Bascom Palmer Ophthalmic Repository with clinical notes from 2014 to 2024. Two glaucoma specialists manually annotated data on glaucoma presence, type, and severity. The models achieved high accuracy in glaucoma diagnosis (94.4% to 97.5%), type classification (94.0% to 97.1%), and severity determination (94.0% to 95.2%). The language models substantially outperformed traditional ICD-10 codes, which achieved only 89.2% accuracy for type classification and 58.5% for severity determination. The results demonstrate that secure cloud-based language models can effectively transform unstructured clinical documentation into structured research-ready data.