A research team developed SIM-VAIL, a clinically validated system for testing the safety of AI chatbots in mental health contexts. The framework simulates users with specific psychiatric vulnerabilities and conducts multi-turn conversations with popular chatbots including Claude, ChatGPT, Gemini, Grok, and Llama. During 810 conversations with 9 chatbots and 30 simulated user profiles, researchers found that concerning behavior was widespread, though less pronounced in newer models. Risk increased over the course of conversations and was highest when the chatbot inadvertently reinforced the psychological mechanisms underlying the user's vulnerability. The SIM-VAIL system offers a scalable framework for mapping mental health risks and provides a foundation for improving AI chatbot safety.