What does it actually take to put 1,000 humanoid robots to work? Hexagon Robotics President Arnaud Robert joins me to explain how the company is moving its Aeon humanoid from pilots and demonstrations into real industrial production, including a plan with Schaeffler that could eventually scale to 1,000 robots.
And the big lesson is that scaling humanoids isn’t about taking one robot doing one task and multiplying it by 1,000. It’s about building a multipurpose fleet that can move between jobs as demand changes across a factory.
Watch here:
We talk about:
- Why humanoids need to be multipurpose to justify themselves
- How Aeon gathers training data while doing real work
- Hexagon’s “sim-to-real-to-sim” training loop
- Why customer-specific factory data matters
- The difference between a successful pilot and a production-ready robot
- The two metrics Hexagon cares about most: cycle time and human intervention
- What Hexagon learned from early sensor-fusion and actuator challenges
- Why Aeon uses wheels instead of legs — and can still climb stairs
- How humanoids could eliminate bottlenecks instead of simply replacing workers
- Schaeffler’s “robot gym” and the train-deploy-scale path to 1,000 humanoids
This episode of NEXT with John Koetsier is sponsored by KindBody Fitness: AI-powered fitness for all the health and none of the gym bro nonsense. Check out KindBody Fitness today.
Robert’s view of the factory of the future isn’t a lights-out facility with no people. It’s a highly autonomous factory where humans and robots work together — with robots handling repetitive tasks and shifting between bottlenecks while skilled workers focus on problem-solving.
This is what humanoid robotics looks like when the conversation moves beyond prototypes and into production.
Transcript: 1,000 humanoid robots in 1 factory
John Koetsier
What does it take to put a thousand humanoids to work in one facility? We’re going to find out. Can’t be easy. We’re chatting with Hexagon Robotics President Arnaud Robert. He’s explained how Aeon is moving from experimental pilots to real factory floors. Hello, Arnaud. How are you doing?
Arnaud Robert
Very good, John. How are you?
John Koetsier
Very, very good. Give us the 30-second background around you. Who are you, and why are you in this role right now? What makes you passionate about this stuff?
Arnaud Robert
Sure. Shortly, I did a PhD in artificial intelligence quite a few years ago. I moved from technology to product and business. I love to be at the intersection of an industry going through a major shift or major innovation push. That’s how I built my career at Disney, at Nike, at Microsoft, and other companies.
And I’ve joined Hexagon Robotics about two years ago because I really think that the movement from automation to autonomy in the industrial space is creating a wave of revolution, frankly, like we haven’t seen before. It’s just exciting to be right in the middle of it.
John Koetsier
Super cool. Talking about things we haven’t seen before, I’ve literally said to multiple robotics CEOs as we’ve been having conversations, I don’t really see the factory that’s going to have 10,000 humanoids in there where they used to have 10,000 humans, because there are different things that they’ll do and there are different levels of automation that we’ll need.
But you are talking about deploying a thousand humanoids. Talk about that. What’s that look like? What’s that feel like? That’s a totally different challenge than putting two in as a test.
Arnaud Robert
Absolutely. First of all, it’s a great opportunity, obviously, to be able to do this. But when you look at a thousand robots, you have to really think about things quite differently. It’s not one robot doing one task that you multiply by a thousand. That would be linear, but not that exciting.
I think it’s really the fact that now you have to think in terms of a fleet of robots and how you would deploy a fleet of robots across multiple factories, not just one, in an optimal way.
And then from there, you kind of go backwards and you say, if you need a fleet of robots, then you need those robots to be able to do more than one thing. Because it’s really about demand, right? Factory floors are quite complicated. There’s demand in a particular area of the factory; it can move to another area of the factory, and so on.
And so we really designed Aeon, our humanoid robot, to be multipurpose. In addition to manipulation, like pick-and-place or moving boxes around, we can do inspection of parts, inspection of space, we can do reality capture. And we think that’s the magic of getting to a thousand robots in a factory: this multipurpose fleet that a customer can deploy based on demand on the floor.
John Koetsier
That’s kind of the thing, right? That’s why you’d have a humanoid. If it’s just doing one job, you basically need an automation solution. It’s always in the same spot. I take this thing, I move this thing there, or I apply this action, I tighten this bolt, I weld this object, whatever.
But if it’s doing multiple things, that’s why you need a flexible system, correct?
Arnaud Robert
Absolutely. And even within one task, I think a humanoid can have an interesting part to play, especially if you need to be both mobile and do a certain task.
As an example, with BMW, we’re doing quite a few different types of manipulation, but one of them is really to move a spoiler from point A to point B, so from one line to another line. And so you need a robot that can, of course, have the dexterity to pick the spoiler up in a certain way, but also transport it a hundred feet away from point A to point B to feed it to the next line, right?
So even within a single task, I think there’s quite a bit of complexity that needs to be solved. But John, as you said, that’s great, but then if the humanoid can do many things, that’s really where the power gets unleashed.
John Koetsier
Talk a little bit more about the sensors that you’ve got on the system, because you talked about training data, and that’s a big deal right now, right? Everybody is scrambling to get all the physical AI training data that they possibly can.
We saw Figure last week formalized and announced they’re doing Index, and it’s spending a billion dollars analyzing that data. I just reported yesterday on another simulated data startup. It’s pretty interesting because you’ve got companies across the spectrum who are using video data, head-mounted video data to learn. You’ve got companies that are using teleoperation rigs, and so the robot is actually doing what the human is telling it to do. You’ve got other gloves with sensors so that people are seeing what hands are doing.
There are so many different options here. You’re saying, hey, we’ve got this platform, it’s going to do the work, and it’s also going to build the data.
Arnaud Robert
Yeah, absolutely. I think you said it quite well. We see three main different types of training data. One is simulation data, as you mentioned. The second one is synthetic data, which is also getting more and more popular. You can auto-generate, especially with generative AI, right? You can generate the videos; you don’t have to record them necessarily. And then you have recording, right? Which could be, as you mentioned, based on teleop, cameras, and so on.
Of course, we use all three of those. Like our competitors, you need to have quite a diverse set of data to be able to train the robots to do things that are quite generalized.
What we found, John, is actually two really interesting things. First, in a specific factory, whether it’s BMW, Schaeffler, or others, the task is actually quite specific. And so you would have to train with that specific data in order for the humanoid to be able to do the task. And that’s data that’s close to the customer, right? They don’t want to see this in the public space.
So we also had to think about how we train a robot using customer data in a way that effectively protects that data but still allows us to train. That was one thing we needed to solve, which we’ve done.
The second one, and especially in the humanoid space, what’s really interesting — and I’m sure, John, you’ve heard about it — is the sim-to-real gap, right? You do something in simulation; when you go to reality, there’s a gap. You’re trying to use other types of training data, whether it’s synthetic or recordings, to bridge that gap.
What we found is actually quite unique. It’s not really sim-to-real, but it’s sim-to-real-to-sim. And it’s that last loop that is really interesting. In that last loop, that’s where our sensors really take flight, if you want. Because with all the sensors we have in the robot — and we can talk about it in more detail — we can actually, while doing a task, record not only the task itself, but the entire environment around the robot.
And it turns out that when you do training, it’s not just about the specific task, because when you go onto a factory floor, many things could happen around that task. A human could be walking by, a tray is in the way, the part was supposed to be on the left, now it’s on the right because somebody placed it there. And so you have to have all that — we call it spatial intelligence, spatial awareness — that you need to have. So our sensors allow us to do this while we’re doing the task.
But more importantly, as we found out, we can continuously record this and feed it back to the training. So as we actually do the task itself, we feed it back into the training. Whether it’s a positive reinforcement, meaning the task was successfully done, or something had changed in the environment, now we can feed it back into the training so that the robot can perform at its best.
So the data space is fascinating. Frankly, it’s evolving very quickly. As you mentioned, certain companies are focused solely on providing those data sets. We think that’s great. But we think the two augmentations I talked about — one, customer-specific data that we need to be careful how we use, and two, the sim-to-real-to-sim, the second aspect being the real-to-sim — we found those two elements are not only useful. We found them to be quite necessary, actually.
John Koetsier
What’s fascinating about what you just said is that where I was going to go with my next question is, it’s always an open question if you’re getting video data: what percentage is usable, right? How many hours of video data do I need to have one good training hour of data?
And it’s not just that you have the video. Is it labeled? Do you know what’s happening there? Do you have a meta-understanding of the task and of the physical realities that are going on there?
Also, if you’re getting training data from the actual robot that’s doing it, you have the whole sensorium of the robot to work with. How hard am I gripping? Where did I move my hand? I was a little bit away. I had to adjust. All that stuff. That’s a very rich training data source.
Sorry to harp on this one thing, but that’s pretty cool.
Arnaud Robert
Yes.
No, it is. And I think you nailed it, right? Is the data useful for training? It’s one thing to just have a very, very complete set of data, but it has to be useful for training.
There are two phenomena that we’re seeing that are quite interesting. One, yes, in general, richer data helps the training. I would say it’s quite intuitive, if you want. But the other thing we’re seeing is also, as algorithms are being developed and learning and training methodologies are evolving, you actually not only need less data, you may need different types of data.
And so what we’re seeing as a really exciting trend for us is that the type of training data is not just about volume, and over time the volume reduces. It’s what type of data is actually useful for the robot.
We have the intuition somehow, right, that we need to train it with the data that we would use as humans to train ourselves. And that’s partially true. But actually, the robot is also having its own layer, if you want, of abstraction. And what we really need is the training data that the robot needs to be able to do the task.
That’s, I think, a bit of a nascent field, I would say, but quite exciting as well.
John Koetsier
Very, very cool. So I keep reporting on a new pilot here, a new pilot there for humanoids. Just one literally this morning in Korea that I saw. All great, all interesting, all that stuff.
What separates a pilot, even a successful one, from a production-ready robot?
Arnaud Robert
Very good question. I think, John, the short answer is that a pilot demonstrates feasibility, but when you go to production, you need two things: repeatable performance and scale.
Usually pilots don’t really address that directly, right? They’re really trying to prove that, yes, you can do the task. You can do it once, you can do it twice. But when you go to production, you have to be able to do it a hundred times in a six-hour shift with a high level of performance.
And you need to be able to scale this, meaning that if one robot can do it, your other hundred robots can also do it, right? So it’s really, to me, from feasibility to repeatable performance and scale. And that transition is not that easy.
For us, we have the privilege of being part of the Hexagon group, which has 34 years of experience in the industrial world. And so we were able to really engage our colleagues from different divisions to understand: what does it really take to go to production?
We’ve kind of built that into the design of the robot, of course, but also into the methodology we use for pilots. So we don’t really do pilots just for feasibility. We already have the seed, if you want, of that repeatable performance and scale.
John Koetsier
Yeah.
Arnaud Robert
And of course, depending on the customer, the pilot addresses that directly or looks at it as a step two. That’s fine, but we need to embed it into our own design of the pilot.
John Koetsier
That’s kind of a cheat code, and I’ve seen that with a couple different companies now. I’ve seen that with Hexagon, I’ve seen that with Agile Robots here where I am in Munich right now, I’ve seen that with Apptronik.
There are companies that are first-principles robotics companies that say, hey, let’s make a humanoid robot, and they started building one. Some of those are doing extremely well and building extremely interesting platforms and great software and everything.
But there’s a difference with companies that have a history in industry and of building robots or automation solutions or other things that are actually in production and have been for a decade or longer. They have an understanding of how things work and how things need to work. That’s quite, quite interesting to me.
When you look at what you’re putting in place, what are the metrics that matter most to you? Is it uptime? Is it speed of the task? We’ve seen there’s a speed gap for most humanoid robots right now versus a human. Is it task success? Number of human interventions? What are some of the key metrics that you look at?
Arnaud Robert
I think the array of metrics is pretty large, but we basically go down to two metrics that really matter. One is cycle time, and one is human intervention.
Cycle time is a great metric, actually, because it combines many different metrics. But at the end of the day — and we talked to literally dozens and dozens of customers to ensure we understood what that really meant for each of them — it’s really about: we have a shift of six hours and, as an example, 500 parts need to be moved in that six-hour shift. That’s our metric.
It really doesn’t matter, in some sense, if the robot can do 400 of those in the first two hours, and then, of course, it gives you more time to do more. Or whether it misses once in a while, but it can recuperate. As long as the cycle time is met, that’s the core metric.
For us, that was a bit insightful, actually, because then you focus way less on specific submetrics and you really focus on performance in the shift. And so you have really interesting compromises that you can do from that point, right?
As you said — and it’s true for all humanoids today — the speed of movement is just slower than a human worker who has been doing it for the last five years and knows exactly the positioning and so on. That will change over time, but currently that’s the case.
But if speed is one thing, then you also look at metrics that affect cycle time, like how many human errors are there in the six-hour shift? And it turns out it’s not zero, right? And so you basically blend all those things, and that’s the metric.
The second one is human intervention. And the reason that’s so critical for us, not only for customers for obvious reasons, but for us as a company, is that we really focused from the get-go that our humanoid Aeon needed to be autonomous.
That’s one of the reasons we have so many sensors. We can have a full view of the environment. We can detect obstacles. We can avoid obstacles, and so on, because the robot really needs to operate in a fully autonomous way.
We don’t believe in a factory you can do teleoperation. We don’t believe in a factory you can be constrained to a single space, because as you pointed out before, then you have other solutions of automation to do this.
And so if you take that into account, then the real metric is human intervention, right? You cannot be autonomous and then also have frequent human intervention. That would, of course, be in conflict. So that’s the second metric we really focus on.
John Koetsier
Cool. As you’ve been working on Aeon and how it works and all those things, what are some of the most instructive failures that you’ve had so far? What was one of the things that happened where you thought, wow, that really taught you something?
Arnaud Robert
We had a few, but I’ll expose just one or two.
I think the first one is that we have a lot of sensors in our robot. So, very different types of cameras. We have shutter cameras, panoramic cameras, we have time-of-flight cameras, and so on.
To have a good understanding of the environment and to be able to perform the task, whether it’s perception, manipulation, and so on, you need to have sensor fusion, right? Meaning we capture all this data, but at any point in time, we need to detect which data is useful for the task in this environment.
Although it’s simple words, it’s actually quite complex when you have so much data coming in. Those are high-resolution cameras. We need to have a lot of algorithms that detect this.
So one of the earlier failures we had — we’ve since, of course, surmounted it — was to be a bit too simplistic about, well, we’re doing manipulation, so let’s just focus on the camera that’s looking at the arms and the object, and that’s what we’re going to do.
Then we realized, wait a minute, we also have to look at the environment. Wait a minute, we also have to look at Aeon itself, where it is in space. Is it well positioned? And so on.
And so we had to, I would say, go from simple algorithms that were a bit matter-of-fact to really complex AI-based algorithms that do sensor fusion.
So again, the failure was that we underestimated quite a bit the complexity, but also, on the flip side, the richness of that sensor fusion, as one example.
The second example, I would say, I wouldn’t maybe call it a failure, but I would call it a challenge, and I think it’s industry-wide, is that the actuators for humanoids are pretty recent when you really think about it.
John Koetsier
Mm-hmm. Mm-hmm.
Arnaud Robert
Even humanoids two years ago, most people were taking off-the-shelf actuators from other robotic types of environments or other environments and then putting them into the form factor of a humanoid. And so you had actuators that were just not designed for the tasks, right?
John Koetsier
Mm-hmm.
Arnaud Robert
So they were either very, very high torque but low speed, and so on. And so that was a big learning for us, which was, wait a minute, the actuators, just like muscles in our arms as humans, actually really matter.
And again, you could say, well, that’s trivial. Of course you would do it by spec. But it’s not just about the spec; it’s how all those actuators work together.
John Koetsier
Mm-hmm.
Arnaud Robert
In a single arm, we have seven actuators, for example, in our Aeon. And so they need to work extremely well together, right? And so I would say that was a big challenge that we had to solve as well.
Other companies, we are very aware of this, are building their own actuators because of that challenge. We decided to still do third-party actuators with partners, and we’re working very closely with them to get actuators with the specifications we need.
Now that we’ve learned quite a bit — Aeon at this point has done literally tens of thousands of tasks — just from that data you learn quite a bit about what you need.
John Koetsier
Mm-hmm.
Very cool. Super interesting stuff. You have a wheeled base, and that seems most popular right now, frankly, for those companies that are building humanoids for factory scenarios, logistics scenarios, those sorts of things.
There’s a stability. I can hold things farther out from my body because of that stability. I’m not using energy to stand because I’m on wheels, all that stuff. And I’m in a flat environment. There’s also a regulatory environment within which those things fit that bipedal humanoids currently don’t have.
Talk about that. Is that something that you’ll stick with forever? Is bipedal something you’re looking at eventually down the line? How do you feel about that?
Arnaud Robert
Yeah, I think the short answer is that we’re very happy with the wheels, and we’re going to stick with it.
The longer answer is that we actually looked at bipedal solutions. We looked at different types of wheel solutions, by the way, not just the ones we have now. We’ve looked at different kinds of morphology of the wheel and so on.
It turns out, John, actually, that the wheels are quite complex. So in terms of power management, bipedal is actually way more robust if you’re standing still to do a task.
John Koetsier
Interesting.
Arnaud Robert
But as we discussed, if you have to stand still, it’s a limited set of tasks. And the wheels really get the full value, if you want, when you also have to have some mobility around the task, whether it’s moving boxes from point A to point B or, as I mentioned in the BMW example, a spoiler from point A to point B over quite a bit of distance.
And distance is not, by the way, a straight line. We have to go around objects and so on. So we’re very happy with the wheels. We built our own power management system, actually, for the wheels for that specific purpose.
And it turns out that wheels allow you to do not only efficient locomotion when you’re moving — that’s clear — but we can also go upstairs, for example. We can go over objects. So we don’t see any limitations of the wheels, which was, for us, the biggest question mark when we went to wheels: will that limit us in any way?
It turns out it doesn’t. And it turns out, actually, that even going upstairs, for example, is very efficient with wheels, because you can use the inertia of the wheel. As it goes into one stair, it starts rolling and it kind of moves the body forward, and then so on.
So it’s been quite nice.
John Koetsier
Nice, nice.
Arnaud Robert
So we’re quite happy with that design. Very few companies have it, by the way. Even in the industrial space, the vast, vast majority have a bipedal solution.
John Koetsier
Very cool. Mm-hmm. Mm-hmm. Interesting. I mean, stairs were the counterexample that lots of people give. Well, that’s why we need a two-legged robot, right? You’ve solved that with the wheels. That’s pretty cool.
Arnaud Robert
Yeah, absolutely. Absolutely.
Yes.
John Koetsier
Talk about the factory of the future. It’s being invented as we speak. You’re one of the companies that’s inventing it. What does it look like? What does it feel like?
We have different visions of that. We have some that are 100% automated. It’s a dark factory. There are no lights on or anything like that, right? Others, there’s a mix.
When you look at the factory of the future or the logistics warehouse facility of the future, whichever one you’re most focused on, what’s that look like and what’s that feel like to you? Who’s there? Who’s not there? What equipment is around?
Arnaud Robert
Yes. I would say the factory of the future for us is one where autonomy and strong human-robot interactions coexist. That would be my one sentence. So very, very high autonomy, but very strong human-robot interaction.
And I didn’t say robot-human interaction; I said human-robot interaction, right?
Autonomy is, I would say, clear, right? What we’re really trying to address — and that, to me, is the core of it — is because, yes, you can invent a future or some projection, but we always like to be grounded in the reality of that future.
And the reality of that future is really today, which is there’s a labor shortage in many, many factories and industries in the U.S. and in Europe.
So you really have a shortage of skills, and a humanoid is a perfect complement to that. At the same time, you have shifting tasks, right? And the one we’re seeing, which is quite interesting that we need to solve and that formulates a bit the factory of the future, is that you need to reduce bottlenecks in a factory. That’s one of the key things they need to do.
And a bottleneck could be because you have only three people that can do inspection.
John Koetsier
Mm-hmm.
Arnaud Robert
One of them is sick, one of them is on vacation, and then you have only one left. So what do you do? We call it complementing the workforce, right? And that’s really critical when you look at the productivity of a plant.
Of course, every single task in a line is important from a productivity standpoint. What we found is equally, if not more important, depending on the industry and the company, is actually avoiding bottlenecks. Because the bottleneck really means that effectively the rest of the factory line waits, right?
John Koetsier
Mm-hmm.
I can do 99% of the job. I can’t give you a completed product because I can’t do the 1%.
Arnaud Robert
Exactly. And so that’s why the factory of the future for us is autonomy. And autonomy is also solving those bottlenecks in an automated way.
As I mentioned, for us it’s multipurpose. So if there’s a bottleneck in inspection, you can send 50 Aeons to get rid of that bottleneck and the factory line is back running.
And the strong human-robot interaction — why we feel it’s very important in that future state, if you want — is that we see the humanoids as really doing repetitive tasks, as I mentioned, trying to complement the workforce when there’s a bottleneck.
But what that really does, actually, it unlocks the humans, who are the ones who have known these processes for years, inside out. And they also have dealt with a lot of the unknowns in the factory, because they happen, right? Things change, et cetera.
And then they can move and shift to problem solving. And so now you have expert workers who know exactly how to do a task, and then suddenly a humanoid can do some of that and they can start focusing on, okay, but now that that’s done, how would I change the production workflow to be more efficient? What else can I do to maybe manage the fleet of robots instead of doing the task myself?
And that’s, for us, a big element of the future. So yeah, I think for us, that’s really the two key elements.
John Koetsier
It’s only when you try to automate a process that you really learn how much you don’t know about a process, how much the humans were just doing because it needed to be done. It wasn’t written up in a manual—
Arnaud Robert
Absolutely.
John Koetsier
—or something like that, or that you didn’t even know that they were getting inputs that were slightly incorrect, but they were accommodating for that because they’re smart and they’ve been working on it.
When you try and automate that, then you’re going to find all those things, and all those edge cases become super important because you can do 97% and not the 3%.
We started this chat by talking about your work with Schaeffler and the plan to get a thousand humanoids operating there. What needs to be done? What needs to happen? What steps do you need to take? What quality bars do you need to surpass in order to make that 100% real?
Arnaud Robert
First of all, we have to be very pragmatic, right? So we’re taking a stepwise approach with them. We’ve announced that we participate in the Schaeffler gym, right? That’s where they basically train their humanoids, which is a very unique setup, actually. It’s an exact replica of the production environment, but where you can train the robot. You can train Aeon.
And so the path to a thousand for us is really: train the robot in the gym, then deploy in production, and then scale that production.
The training in the gym does two things, of course. If you train well, then when you deploy in production, it should work as is. But it’s the scaling aspect that’s also quite interesting in the gym. If you do it very, very well, then you can not only train one Aeon, you can, say, train a hundred Aeons, a thousand Aeons, through basically replicating what you’ve learned on one to all the other ones.
And so that’s kind of the path towards the thousand, if you want.
The second dimension, I would say, is we’re looking at multiple use cases in that gym, right? So it’s not just one task that Aeon needs to do. We’re looking at a number of tasks, and some of them we’ll find out that Aeon is particularly good at, and that’s a perfect scaling opportunity.
John Koetsier
Mm-hmm.
Arnaud Robert
Others, we may find that in today’s technology it’s still a bit difficult to realize the task at the same cycle time as a human would do.
And that’s fine, right? You learn through that experience. And then you basically deploy the use cases that provide you the most scale.
And of course, over time, we will address more complex use cases in the gym. And then, same thing: we train, we deploy, we scale.
So the train-deploy-scale approach for us works quite well. We’re doing it also at other customers, by the way, but that’s the key for us with Schaeffler.
John Koetsier
Super cool. Thank you so much for taking this time.
Arnaud Robert
Absolutely. It was a pleasure, John. Thanks for having me.
