AI Quality Testing Essentials (ATE)

AI Quality Testing Essentials (ATE)

Set up your first eval and get it running in 4 hours. Discover, by doing, why AI Quality Testing is a discipline in its own right.

Flash Workshop
4 hours
Upz!

Cohort completed or already started.

Leave us your details and we'll let you know when we open the next cohort.

What is it about?

This is a hands-on workshop designed for product and testing professionals looking to gain their first practical experience in evaluating AI products. Over the course of four hours, you will work on a real-world case study, configure actual evaluations on a live platform, and witness firsthand how the traditional testing mindset transforms when applied to a probabilistic system. We will focus on a simple, well-defined scenario where verifiable criteria and subjective criteria coexist within the same problem.


You will leave having completed your first evaluations, with a concrete understanding of how much more there is to learn, and with the insight needed to start engaging in more informed conversations within your team.


We will be working with OpenAI Evals, currently the industry-standard platform for AI Quality. No programming knowledge is required, as the majority of the work is performed directly through the user interface.


It is necessary to have an active OpenAI account and at least USD 10 in available credits to complete the hands-on exercises.


Tools

We will use the following tool

OpenAI Evals

Who is this for?

Profesionales de Producto

(Product Manager, Product Owner, Product Lead) que trabajan o van a trabajar con productos basados en IA y quieren empezar a incorporar criterio cuantitativo para validar calidad.

Profesionales de Testing

(QA Engineer, Tester, Quality Lead) que ven cómo la IA está cambiando la disciplina y quieren empezar a explorar, con las manos, qué cambia y qué se mantiene.

After completing this workshop you will be able to:

Construir tu propio dataset sintético: pensar las dimensiones del problema, escribir el prompt de generación, y producir 15-20 tickets sintéticos listos para correr la primera evaluación.

Leer outputs de un LLM de forma sistemática y construir una primera taxonomía de modos de falla.

Configurar y correr una evaluación determinística sobre criterios verificables como formato o estructura.

Construir un LLM-as-a-judge de cero: escribir la rúbrica para tu caso, traducirla a un prompt de juez, configurarlo en la plataforma y correrlo sobre los outputs del clasificador.

Comparar los veredictos de tu juez con tu lectura humana y descubrir, en vivo, por qué un juez consistente puede equivocarse sistemáticamente, y por qué eso abre la pregunta de la alineación.

Salir con un eval funcionando que puedes llevarte y adaptar a tu propio caso.

Your Instructor

Martin Alaimo

Martin Alaimo

Since 2009, he has worked with more than 200 organizations and supported over 8,000 professionals in their career development journeys.

His approach is situational and hands-on, delivering immersive learning through innovative experiences that enable practical, immediately applicable outcomes—especially in areas often overlooked by traditional academia.

He has spoken at more than 30 conferences across the United States and 14 countries in Latin America and Europe, and is the author of six books on product and digital innovation.

His most recent book, AI Strategy Workshop, provides tools to move beyond the “feature factory” mindset and integrate artificial intelligence with strategic intent and real business impact.

He is the founder of Verica, a platform for running evaluations (Evals) on LLM-based products, designed for organizations that need to measure the quality of the outputs generated by their AI products.

As part of his commitment to innovation, he is an organizing member of Product Tank, the world’s largest Product Management community.

He is one of the few experts to hold the highest-level certifications in Agile practices: Certified Scrum Trainer (CST), Certified Enterprise Coach (CEC), Certified Team Coach (CTC), Certified Agile Leadership Educator (CAL Educator), and Path to CSP Educator.

Visit his complete professional profile and thought leadership activities on LinkedIn.

Ready to step forward?

You're not starting from zero. You're choosing to advance with purpose.

Learning better is also a decision.