The Art Behind Training Gen AI Models: An Interview with Leonardo.Ai’s Lead AI Researcher & Co-founder, Ethan Smith

Insights | Published on

5 min
The Art Behind Training Gen AI Models: An Interview with Leonardo.Ai’s Lead AI Researcher & Co-founder, Ethan Smith

Following the launch of our latest image generation model, Lucid Origin, and in light of its recent entry into the top 10 on the community-driven leaderboard at Artificial Analysis, we sat down with Ethan Smith – Leonardo’s Lead AI Researcher and Co-founder – to unpack how we build and train new models. In this deep dive, Ethan shares the philosophy, data strategy, and hybrid thinking that set Leonardo.Ai apart in the world of generative AI.

The Leonardo Philosophy: Art Meets Science

Leonardo.Ai was built on a belief that technical progress alone isn’t enough. Real innovation happens when engineering precision meets creative intuition. That means placing as much emphasis on dataset curation, aesthetic sensibility, and iteration as on raw compute and performance benchmarks.

Ethan, you spearheaded the development of Leonardo’s first foundational model, Phoenix, and have since contributed to several other projects at the company. How did your experience with Phoenix influence Leonardo’s approach to model development? In such a fast-moving field, what does Leonardo prioritize?

ES:
Phoenix really set the blueprint for how we approach generative AI at Leonardo. It wasn’t just about pushing technical performance – though that was a major focus – we also wanted to elevate aesthetic quality. We built it entirely in-house, backed by significant proprietary research and compute power, and at the time it combined precision, usability, and efficiency in a way the market hadn’t seen before. But what also made Phoenix stand out was the emphasis we placed on aesthetics, just as much as technical excellence. That balance between technical sophistication and creative appeal has guided every model we’ve built or worked on since.

How do we approach building a training models differently to some of the big names in the market?

ES:
I genuinely think one of the biggest differentiators in the world of visual AI right now, and for us, is taste. A lot of the improvements you see in models today don’t come from radically new architecture – they come from scale and knowing how to wrangle your data better than anyone else. Curating a dataset well is everything right now. Scaling compute matters, yes – but we’ve gotten so far by pairing this with highly curated datasets, guided by good taste. I’d say that’s a core part of it.

When it comes to data wrangling and curating datasets, how manual is the process behind the scenes?

ES:
Very manual. We strive to automate where we can, because there’s just so much of it and it’s a lot of work. But even in the best-case scenarios, researchers are still sifting through data manually. Looking at data and having to determine whether it’s aesthetic or not. It’s a lot of work.

Is that just because there’s no way to automate “taste”? You still need human judgment?

ES:
Exactly. We try to automate it because of sheer scale – but at the end of the day, no algorithm can yet fully replicate a human eye for what’s aesthetically strong. Personally, I don’t have that kind of visual taste. I’d love to reduce it to a science. But that’s why this kind of curatorial input is so rare in AI research – it doesn’t fit neatly into a numerical or rigid framework.

Why do you think it’s so rare to see taste combined or mentioned in the AI research side of things? Why is it a rare combination?

ES:
I think it’s hard to pull off, and requires a lot of balance. And I’m seeing that it tends to the be the smaller startups who are nailing the taste side of things better than the larger organizations, typically.

I think it’s very common to have types who think like me – who want things to be down to a rigid science and a numerical or quantitative way. But that’s why startups like us succeed here. Because we bring in diverse talent – people who are technically strong and also have a deep understanding of aesthetics. That hybrid skill set is critical.

So the competitive edge now lies not just in technical precision, but being able to successfully marry that with strong artistic curation?

ES:
In a way. The early breakthroughs were about performance – prompt adherence, anatomy, speed. Now, gains are mostly driven by good data and smart training – but that whole process is filled with nuance. Selecting the right data is key, and its started to become an art. Data refinement is kind of inherently taste.

How do you define what makes something look ‘good’ or ‘bad’ in terms of taste? Is it based on some objective criteria, or is it just learned from training on a large number of images?

ES:
It’s very philosophical. Like I said, I’m not someone who has it, but there are those in our team who do. What we’ve done is we’ve created these data sets, we’ve trained the model, and then we go to the tastemakers in the team, get their input, and then we go back to the drawing board to see if it worked or not. It’s a constant iteration of our data sets and trying out different training hyper-parameters. Different training configurations.

Because it’s not just about the dataset itself. It’s also about how the data is weighted and balanced within the model, which can lead to very different outcomes?

ES:
Right. It’s a bit voodoo in that you can’t really know for certain exactly how some of your decisions impact things in a very rigorous, scientific way. You kind of tweak things here and there. It approaches a bit of an art.

You want to put a model through the ringer, train it on all these robust scientific principles, and then have it come out with a clear result. But ultimately, when the model comes out, you kind of have to just evaluate it subjectively yourself a little bit, especially important for image models, rather than just depend on strict evaluations.

How do you see the future of generative AI evolving, and what will it take for generative AI models to truly be considered successful?

ES:
Good taste in AI means building with intention – not just precision. It’s about knowing what looks good, feels right, and will resonate with creators. This mindset influences everything from how training runs are configured to how outputs are evaluated. It’s not just about chasing benchmarks. It’s about building models people want to use – because the images feel right. All the while, maintaining rigorous technical excellence is imperative. Advancing cutting-edge performance remains foundational, serving as the critical enabler for unlocking the full creative potential of generative AI.

The future of generative AI belongs to the teams who can bring technical excellence and taste together – because that’s where models stop being just impressive, and start genuinely inspiring and resonating with creators.