Clinical Psychology Expertise for AI Evaluation, Safety & Quality

20+ Years Clinical Experience | AI Safety & Evaluation | Behavioral Health

AI Safety & Behavioral Health

Clinical Psychology Expertise for AI Evaluation, Safety & Quality

Dr. Steve Orma, Psy.D.
Licensed Clinical Psychologist | AI Safety & Evaluation | Behavioral Health

Artificial intelligence is increasingly being used to provide information and support related to mental health and other health concerns. Ensuring that these systems are accurate, safe, nuanced, and genuinely helpful requires more than technical expertise.

It requires clinical judgment.

I am a licensed clinical psychologist with 20+ years of clinical experience and extensive expertise in behavioral health, including anxiety, insomnia, eating disorders, depression, stress, relationships, and cognitive behavioral therapy.

I also have experience evaluating AI-generated health information for accuracy, quality, and safety, bringing a clinician's perspective to the development and improvement of AI systems.

Bringing Clinical Judgment to AI

AI systems can produce responses that appear convincing while containing subtle inaccuracies, inappropriate recommendations, missing context, or potentially harmful guidance.

Clinical evaluation requires asking questions such as:

  • Is the information clinically accurate and appropriate for the specific context?

  • Does the response recognize important risk factors, warning signs, or limitations?

  • Could the response unintentionally reinforce unhealthy thoughts, behaviors, or misconceptions?

  • Does it appropriately communicate uncertainty rather than presenting questionable information with undue confidence?

  • Is the response clear, empathetic, and genuinely useful to the person receiving it?

  • Does it appropriately balance helpfulness with safety, including recognizing when professional evaluation may be warranted?

These are areas where real-world clinical experience can provide an important layer of evaluation beyond factual accuracy alone.

AI Safety & Evaluation Experience

I currently contribute my clinical and behavioral-health expertise to work involving the evaluation of AI-generated health information.

My work includes assessing AI outputs for factors such as:

Clinical accuracy
Evaluating whether health-related information is consistent with established clinical knowledge and appropriate for the context.

Safety
Identifying responses that could potentially cause harm, provide inappropriate guidance, overlook important risks, or fail to recognize when additional professional support may be warranted.

Quality & usefulness
Assessing whether an answer actually addresses the user's underlying question and provides clear, relevant, practical information.

Clinical nuance
Recognizing situations in which technically correct information may nevertheless be inappropriate, incomplete, misleading, or poorly framed for a particular user.

Behavioral-health reasoning
Applying knowledge of human behavior, cognition, emotion, motivation, and psychological processes when evaluating AI responses.

Structured evaluation & feedback
Clearly identifying problems in AI-generated responses and communicating the reasoning needed to improve system performance.

Because some of my AI work is subject to confidentiality obligations, I do not disclose proprietary project details. I can, however, discuss my areas of expertise and the types of AI evaluation work I perform.

Why Clinical Experience Matters

A response can be factually correct and still be clinically poor.

Consider a person asking an AI system about anxiety, insomnia, eating disorders, depression, or a frightening physical symptom.

A useful evaluation cannot stop at:

“Is this statement technically true?”

It must also consider:

“Is this the right information for this person, in this situation, presented in a way that is accurate, safe, and genuinely helpful?”

After more than two decades of working directly with people, I have developed extensive experience recognizing the difference between information that is merely correct and information that is clinically appropriate and useful.

That perspective can be valuable when evaluating AI systems operating in complex human situations.

Areas of Behavioral-Health Expertise

My clinical background includes extensive experience with:

  • Clinical & Behavioral

    • Anxiety

    • Depression

    • Stress

    • Behavioral change

    Psychological Expertise

    • CBT

    • Test anxiety

    • Relationships

    • Self-esteem

    Specialized Experience

    • Insomnia

    • CBT-I

    • Eating disorders

    • Work-related issues

    • Adult mental health

I have also developed particular expertise in cognitive behavioral therapy for insomnia (CBT-I) and have spent more than a decade helping people overcome chronic insomnia.

My clinical work has given me extensive experience translating complex psychological concepts into language that people can understand and apply.

What I Can Bring to an AI Team

I can contribute a combination of clinical expertise, behavioral-health knowledge, analytical judgment, and experience evaluating AI-generated health information.

Potential areas of contribution include:

AI Safety & Evaluation

Evaluation of AI-generated health and behavioral-health responses for accuracy, safety, appropriateness, and quality.

Behavioral-Health Expertise

Clinical perspective on anxiety, depression, eating disorders, insomnia, stress, cognition, behavior, and other psychological concerns.

Human-Centered Evaluation

Assessment of whether AI responses are understandable, appropriate, empathetic, and useful from the perspective of the person receiving them.

Quality & Risk Assessment

Identification of subtle problems that may not be apparent from a purely factual or technical evaluation.

Expert Review & Feedback

Providing clear, structured, clinically informed feedback to help improve AI-generated responses and evaluation criteria.

Content & Knowledge Evaluation

Reviewing health-related content for clinical accuracy, nuance, clarity, and appropriateness.

Media & Professional Recognition

My clinical expertise has been featured in media outlets including TIME, NPR, Forbes, Reader's Digest, Shape, Women's Health, Huffington Post, and Men's Health, and I have served as a consultant and contributor to the Calm app, including its Sleep Stories content on sleep science

Interested in Working Together?

I am interested in consulting and expert opportunities involving:

  • AI safety

  • AI evaluation

  • Health AI

  • Mental-health AI

  • Behavioral-health AI

  • AI quality evaluation

  • Human-centered AI

  • AI training and evaluation

  • Clinical content evaluation

  • Health-information quality and safety

If you're developing or evaluating AI systems that interact with health or behavioral-health information, I'd be interested in discussing how my clinical expertise could contribute.

Get in Touch → steveorma@gmail.com