I met Ivan around the launch of Pleias and was immediately struck by his direct answers, positive energy, boldness, humor, and humility, all while being clearly exceptional in his field. He is physically impressive, but his demeanor should fool no one: he’s one of the kindest people you can meet, and if he may come across as strongly opinionated, he’s actually always ready to listen and change his mind when confronted with a solid argument. There’s no ego, only a drive for truth and improvement. Even better: you can feel that Ivan has the grit to pass on knowledge like few others. It’s pure enjoyment for him to help others progress, which to me shows he has all the qualities of both an entrepreneur and a professor. And indeed, he is both.
The Journey of Ivan Yamshchikov
As always in The Human Layer, could you start by introducing yourself, specifically as someone working in the technical field?
My name is Ivan Yamshchikov. I am a research professor at THWS, a university of applied sciences in Würzburg, where I lead the CAIRO AI research institute and work on various aspects of generative AI. I am also the co-founder and CPO of Pleias, a French AI lab rooted in open-science principles that builds distinctive LLM-powered solutions to accelerate enterprise AI adoption.
Can you tell us more about Pleias, your current venture with Anastasia Stasenko, and Pierre-Carl Langlais?

Right around Christmas 2024, Anastasia called me and asked what my take on the big LLM players was. I told her that: “It would be cool to build an LLM company that is not based on stolen data”. To which she replied that she wanted to build exactly that company, and I said I was in. In four months, we shipped Common Corpus, the largest and most ethical data collection for LLM pre-training.
We won several European HPC grants that allowed us to build our first generation of LLMs trained on fully open data. We then demonstrated that small reasoners can provide reliable reasoning sources. This work solidified our long-term vision: we believe that a big portion of AI systems will not be an omnipotent behemoth sitting somewhere on the cloud, but rather swarms of small, efficient, highly specialized models deployed differently depending on the use case.
This vision is what defines our current work. Our first product based on our experience is Pleias Stratum, a data layer that simplifies AI adoption for highly regulated industries. Most AI pilots fail, since real data is raw, siloed and messy.
Stratum allows a CDO to focus their team on the tasks that create business value, while our solutions handle formatting, chunking, personal-data anonymisation and synthetic-data harmonization, making the data easily accessible for any AI application that follows.
This week we released SYNTH, a synthetic dataset for model pretraining, along with two models that demonstrate our know-how. Well-cooked synthetic data, used correctly during pretraining, allows us to build very small models that beat the SoTA on popular benchmarks.
You were born in Russia in the 1980s. What was your first contact with computers?
My family couldn’t afford a computer til I was 12, so my most vivid early memory is the one my dad had at the office where he worked. He often had extra hours on weekends when I was in primary school, and sometimes I could go with him and spend several hours playing Heroes of Might and Magic. I’m still not sure I can imagine a better weekend.

You describe yourself as a “radical techno-optimist.” Can you explain what you mean by that?
To me, a radical is someone who cannot compromise on certain beliefs. For me, one of those beliefs is that technology has a net positive effect on the human condition. This is why I call myself a radical techno-optimist.
In the last ten years, I have had thousands of discussions with different kinds of techno-doomers. I gathered many arguments for why technology is good. The main conclusion from these conversations is that techno-pessimism is an axiom for most techno-doomers. They simply believe that technology is evil and build all further reasoning on top of that belief. No matter how much data or historical context I bring, they will not change their mind.
They see the claim that technology is bad as self-evident and not something that requires proof. This means that instead of trying to convince them that one of their core beliefs is wrong, I can have a more constructive discussion by treating their position as an alternative belief system and acknowledging it even if I respectfully disagree.
So now, before engaging in a deeper discussion about the future of humans and tech, I tend to ask several basic questions, like “is there less poverty on the planet now than in the 19th century?” or “do we live longer, safer and more fulfilling lives now than 100 years ago?”
The answers to these questions are very obvious, you can look up the stats. So, if you think that our conditions deteriorated I can immediately understand that this is your axiomatic belief. You do not need arguments or stats, you just accept this as a universal truth. The interesting thing is that we still can have a productive conversation. If anything, stating your axiomatic beliefs explicitly, actually helps to have constructive discussions.
So when people ask me about my view of the future, my expectations or my forecasts, I state clearly that I am a radical techno-optimist and that my predictions rest on a fundamental belief that technology is a net positive. If you want to hear what I have to say, you’re welcome. If this contradicts your fundamental beliefs, I’m not your guy.
You are still a researcher and a university professor as head of AI institute, and now also co-founder of an AI startup building a foundation model. Could you describe the differences between both roles, as well as the overlap and benefits?
At CAIRO, we build our AI master’s program around one insight: great researchers and great founders share several foundational traits. Both must be intrinsically motivated, comfortable with long periods of uncertainty, capable of first-principles reasoning, and able to work across disciplines. Modern science and modern startups both reward people who can think rigorously, collaborate effectively, and push through years of ambiguity before seeing real results.
So, in my opinion, the differences between the two paths are mainly structural. Research allows longer curiosity-driven detours, while entrepreneurship is constrained by market validation and limited runway. Yet academia is also a market of ideas, and topics with high societal demand create faster career trajectories, just as markets shape startup success. The feedback loops are different in speed, not in nature.
That’s why our master’s program is designed for people who haven’t yet chosen between research and entrepreneurship. We train a shared mindset, people who can explore deeply like researchers and execute pragmatically like founders. I believe that these hybrid profiles are exactly the ones driving today’s most meaningful innovation.
Innovation With Openness
Pleias takes a pure approach to open-source LLMs, both fully transparent and ethical. The first wave of LLMs marked the triumph of closed models trained on controversial datasets in pursuit of performance. Now, open-weight models are catching up quickly. Yet, none of these are open in the pure sense that Pleias is. What is your take on this trend, and do you think this could go as far as giving access to the datasets like you do?
People often confuse what “open” means in AI today. Most so-called open-source models are actually open-weights models, where you get the trained network but not the data or the training loop. That is valuable in itself. I need to acknowledge that without open weights, most modern LLM innovation simply would not exist. However, true scientific progress requires fully open, reproducible work, including data, code and training. That principle is centuries old, and it applies to AI as much as to any other field.
We started our company two years ago around Christmas with a simple question: can we build an LM lab that is truly open end-to-end. We released Common Corpus several months later, trained models on it and saw them outperform expectations, sometimes matching models three times their size. That momentum led to our SYNTH dataset, and then to two SOTA-for-size LLMs (Pleias Baguettotron and Pleias Monad) trained on it.

Along the way, we accumulated the real competitive edge, which is not diminished by sharing data and code. In our opinion, sharing strengthens the whole ecosystem and attracts talent without eroding our advantage.
We could not build Stratum if we had not shipped all those fully open assets first. The core thing we learned is that it is not enough to train SLMs to build competitive AI products. However, you cannot build amazing AI products if you do not train your own models from scratch.
Open-weight models allow companies to fine-tune and to keep their data within their own perimeter. From your perspective, what reasons should motivate companies to switch to a fully open-source model?
A fully open model matters because you can actually audit what it was trained on. With rising concerns around model poisoning and AI-driven security breaches, closed-data training pipelines create real blind spots. If you don’t know the data, you can’t guarantee safety. Openness makes the entire training process verifiable and reduces systemic risk.
That said, we don’t believe fully open models are a universal solution. In practice, more than half of AI pilots fail, open or closed, not because of the model but because the data inside the company isn’t ready. It is messy, inconsistent and not AI-native. This is where we saw a gap in the market: companies need something that accelerates time-to-value, something a CDO can hand to their data science team so they can focus on use cases rather than data janitorial work.
Pleias is a bootstrapped lab, yet manages to ship impressive small models and innovate. SYNTH, released very recently, is a great example with fully synthetic data, carefully crafted, delivering strong performance with fewer than 300 million training tokens compared to the trillions used by others. What is the genesis behind this project?
Once we realized that data readiness is the biggest bottleneck in AI adoption, it became obvious that we needed something that showed just how dramatic the gap really is. That is why we built SYNTH. It was our way of proving, with real results, that when you curate data properly and use synthetic augmentation, you can create an AI-ready data layer that unlocks performance far beyond what raw, messy enterprise data can support today.
Approximately a year ago, we discussed the idea of the “minimal model” with our CTO Pierre-Carl Langlais. He had an intuition that properly shaped data could enable unprecedented model performance. So SYNTH became an opportunity to validate his gut feeling.
We published two models trained on it. Monad, with 56M parameters, turned out to be the smallest viable model to date. Baguettotron, with 321M parameters, is state-of-the-art among models of its class on various benchmarks (MMLU, gsm8k, HotPotQA), even though we used orders of magnitude fewer tokens to train it.

Working on SYNTH forced us to think differently about what “training data” means. Instead of collecting diverse internet text and hoping the model would learn everything, we decided to engineer specific capabilities: semantic bridging between concepts, query expansion, multilingual harmonization and constraint-based reasoning. We wanted to see how a synthetic pipeline could create shaped data. Data designed to instill particular transformations, particular ways of connecting information, particular reasoning patterns.
The real lesson we took from synthetic data efficiency is not just “you can train smaller models for a very low cost” but also “context preparation is as important as the model itself”. We hope this release shifts the current majoritarian paradigm for deploying AI, since it is extraordinarily inefficient and prone to failure.

You have mentioned plans to adapt your dataset for domain-specific use cases such as legal, medical, and technical documentation. But do you think this approach could also extend to much larger models with general-purpose applications?
In technology, we know the classic distinction between extensive and intensive growth. Extensive growth means scaling output by simply adding more resources, more land, more mines, more parameters. That is essentially the big-model strategy: increase size and hope performance follows. It works for a while, but it is fundamentally a resource game.
Intensive growth is different. It is what happens when you take the same field and introduce tractors, irrigation or better seeds. You get more yield from the same resources. And historically, every major technological revolution, from agriculture to industry to computing, has come from intensive improvements, not from scaling brute force indefinitely.
That is why we do not worry about the players who are all-in on parameter inflation. It is an extensive-growth bet with ruthless unit economics. When the music stops, those extensive business models tend to pop loudly. The long arc of history favors the teams that learn how to extract far more value from the same compute, the same data, the same substrate. And we are building for that future.
You switched from a CSO to a CPO position a few weeks ago. How excited are you to explore this new discipline?
Research is fun, but when you see a product–market fit, you go for the jugular.
You published the Common Corpus a year ago, the largest public domain dataset. Beyond the transparency aspect part, have you seen others start building with it?

We already know of at least eight labs using Common Corpus for their pre-training, including recent SLMs from NVIDIA, and that is insanely cool. It was trending on Hugging Face alongside a pre-training dataset by Microsoft and HF’s own pre-training dataset, not bad for a small, bootstrapped lab.
Seeing others build on what we ship is awesome!
Indeed! It has been over a year since you launched Pleias. What has been the real impact of being fully transparent and open-source compared to other labs?
In our experience, it’s very hard to raise money for open source, but that is not only a curse but also a blessing. The level of freedom that we enjoy now is unprecedented because we had to play the start-up game on the “nightmare” level from the start.
Rethinking Scale in Enterprise AI
More generally, you have been a strong advocate for SLMs since the beginning of Pleias. I’ve seen so many praising this as well, but when it comes to usage, it feels people tend to favor big models. When do you think it becomes relevant, or even a necessity to switch to a smaller model?
I see a very different picture here. We probably just talk to different people in different segments of the market. Most of the people I know use some closed solutions for rapid prototyping. Once the pilot works, the next step is often optimization and an orchestration pipeline with a mixture of different models, both big and small.
The problem is that very often the pilot does not deliver the return on investment that the company expects. In those cases the model is not the problem, it is usually the data. You want your data to be ready for AI-first processes and solutions, but you often have to deal with a zoo of sources and formats and a lot of governance requirements on top.
This is where SLMs for data curation and harmonization become useful.
Also, the fact that general-purpose LLMs have proven far more versatile than expected but still do not deliver the expected performance (the Toolathlon agent benchmark shows only a 38 percent success rate for the best model, for instance) makes it sound like a good case for a return to specialisation, don’t you think?
Frankly, I do not think specialisation left us. It is just that if you have enough compute and are under constant pressure from investors or shareholders, you become a hostage of high expectations.
If you have billions at your disposal, a press release that says “we have made a tiny coding model that is actually decent” will simply drown in all the noise. If you promised AGI, a lot of people want some progress towards AGI from you.
So even if you could release a small specialised model, you would rather reuse the data that you could use for it in a bigger, more ambitious project so that your next frontier model scores several extra points on some benchmark.
Do not blame the player, blame the game.
Looking back, what has been the most impressive technological breakthrough since your first steps in tech. The one that genuinely blew your mind?
To me that is CRISPR-Cas9. I think the frontier of our understanding of biology is moving astonishingly fast, and computation plays a huge role in it. If you compare how human life changed between 1875 and 1925, you will see a giant leap across many aspects of living, including electricity, planes, engines and manufacturing.
However, if you compare 1975 and 2025, then computers are almost the one and only area where we see astonishing change and progress. I really hope that in the coming 50 years this dynamic will once again appear in biology and material science.
And to end on a lighter note, could you be our Nostradamus of the day and share some bold predictions for the future of AI across all fields?
We either solve PDF parsing or we are doomed to stay in the cradle of our planet.
Resources:
-
Common Corpus, The Largest Collection of Ethical Data for LLM Pre-Training on Hugging Face, ArXiv