From NeurIPS to the Real World: Key Takeaways from Competing in Pommerman

In 2018, I entered the Pommerman competition at NeurIPS with Chao, Bilal, and Matt, my supervisor. Our agent finished second in the Learning Agents category.
Looking back, what stayed with me was how much the experience resembled applied ML in industry. We had to decide what to simplify, where to add complexity, and which experiments were worth the time before the deadline. Those decisions are what I want to reflect on here.
First, let’s start with a short introduction to Pommerman.
What is Pommerman?
“Pommerman is a multi-agent environment based on the classic console game Bomberman. Pommerman consists of a set of scenarios, each having at least four players and containing both cooperative and competitive aspects. We believe that success in Pommerman will require a diverse set of tools and methods, including planning, opponent/teammate modeling, game theory, and communication, and consequently can serve well as a multi-agent benchmark."1
A Screenshot of Pommerman, agents can place bombs that explode after some timesteps.
You can check this video, it’s quite easy to understand the game after seeing it.
The game resembles a console game, and the simulator allows you to control some of the four players (agents). Some game rules include:
- Each agent can execute one of six actions per timestep: do nothing, move in one of four cardinal directions, or drop a bomb.
- The board consists of passages (walkable), walls (indestructible), or wood (destructible with bombs).
- Maps are randomly generated.
- Bombs explode after 10 timesteps, destroying wood and any agents within their blast radius.
- There are power-ups for players.
- Each game episode lasts up to 800 timesteps.
- There is only a binary reward at the end of the game, +1 or -1.
- The agents only see a region around them, not the entire board (partially observable).
If you want to learn more about how we addressed the technicnal challenges you can check here.
Having a background on the competition, we can go into the lessons learned.
1. Working with Deadlines

When I learned about the competition, I had a conversation with my supervisor along the lines of:
S: "Have you seen this Pommerman competition?"
Pablo: "Yes, sounds interesting"
S: "Yeah, we should submit something"
Pablo: "Ok... but we only have like 3 months"
S: "Great! Plenty of time.
During my Ph.D., I had similar comments from previous supervisors. Conference deadlines are intense, meaning you have to work under pressure. I also thought it meant sacrificing quality.
Years later, my attitude towards deadlines has changed. I’ve even seen extreme cases of having just one week to produce something (hackathon style?).
Does it make sense to have such tight (impossible) deadlines? Maybe.
Now, I try to make sense with the following arguments:
-
Setting a high bar by your manager sometimes means they believe in you. “When our limits are pushed in a healthy, empowering way, it can help us fulfill our potential and provide us with the drive and motivation to excel."
-
Even if your manager doesn’t believe in you, deadlines force you to take action, make decisions, have ownership, collaborate, and produce something tangible.
I also think that once the deadline has passed, there should be some time to take a step back and analyze the outcome.
2. Dealing with Constraints

“Frugality drives innovation, just like other constraints do. One of the only ways to get out of a tight box is to invent your way out.” — Jeff Bezos
Pommerman had many distinct technical challenges: multi-agent, planning, learning, and communication, among others.
Trying to solve only one of the issues wouldn’t have resulted in a good agent for the competition.
-
One example to understand the challenges was the “lazy agent” problem2. What we experienced was that when training a team of two agents, one being stronger than the other, the weaker agent tended to just avoid any actions (it was just sitting duck). The reason was that if the agent started exploring, it would probably get blasted by its own actions, affecting the reward of the entire team.
-
Another example was the time limit for taking an action by the agents of 100 miliseconds.
“If an agent does not respond in an appropriate time limit for our competition constraints (100ms), then we will automatically issue them the Stop action and, if appropriate, have them send out the message (0, 0). This timeout is an aspect of the competition and not native to the game itself."1
Having multiple constraints is no different from “real-world” problems. You will never work on an “easy” problem. You will never start with perfect data. You won’t have enough time to try all the models you want. Your model/agent needs to run within certain limits (e.g., time or memory). Your training data is never complete. Inherently, you are always working with deadlines. Learning to navigate these limitations is crucial for a successful outcome.
3. Focusing on Solving Problems

NeurIPS is about research papers, but competitions try to foster new ideas based on solving problems. It’s a subtle difference, but one that’s closer to how industry works.
One should be careful not to lose focus on the main goal. I’ve seen two examples of how focus might dissipate:
-
Getting Sidetracked by Too Many Experiments. When you start working on a new problem, the focus is there, it’s clear. But after becoming comfortable with the problem, I tend to want more and more experiments. It’s essential to maintain focus and avoid getting sidetracked by too many experiments. Balance is key —push towards a complete solution without losing sight of the primary objective.
-
Losing Sight of the Main Metric. he second problem appears during evaluations and the metrics that matter. For example, in the Pommerman case, we modified the reward function of the agents and had plots of how the reward improved during training. However, the metric that really mattered was the “win rate” (defeating the the other team). In industry, this is similar —you have your model metrics, but there’s also a business metric. One needs to be careful not to forget that the business metric is what really matters.
4. Understanding Trade-offs

“People who are bred, selected, and compensated to find complicated solutions do not have an incentive to implement simplified ones.” ― Nassim Taleb
Two decisions from the competition stayed with me because they involved different trade-offs: we simplified how the agent remembered the board, but added complexity to its training opponents.
Sometimes an approximation is enough.
Our agent could only see part of the board. Information from previously seen regions could still be useful, so we needed a way to remember it. A recurrent network was one option, but it complicated training.
We used a simpler approach: keep the last observation of a region until a new observation replaced it. That information could become stale as the board changed out of view. Still, the approximation gave us good results and only required a simple vector of data, without adding another network component.
Other times, the simple approach leaves too much out.
Choosing training opponents was a different problem. We started with static agents so our agents could learn that blasting an opponent was how to win. That was a useful starting point, but it would not prepare them for the competition.
Here, adding complexity helped. Chao developed an unbiased smart agent to train against, and we saw a large improvement in performance.
These choices made the trade-offs concrete for me. With memory, we accepted potentially stale information in exchange for a simpler implementation and training process. With training opponents, we added complexity because the simpler approach left out behavior our agents needed to learn.
The question I took from this was: what are we gaining, what are we giving up, and how does that affect the result we care about?
5. Taking Ownership

“If all you need to do is what you are told, then you don’t need to understand your craft. However, as your ability to make decisions increases, then you need intimate technical knowledge on which to base those decisions.” ― L. David Marquet
Every member of the team had a different expertise, but we all knew the problem and the goal we were working towards.
- We were independent in the sense that individually we had to think, implement, and validate our ideas in our specific area of expertise. We had ownership.
- We worked as a team because we cared about the result and reaching the goal. We embraced shared learnings —both things that worked and things that did not— which helped to boost trust among the team.
In the end, we took the lessons from everyone and put them together in the agent that was submitted to the competition.
6. Building from Scratch

“Take a simple idea and take it seriously.” — Charlie Munger
Working on the competition was a clear, simple idea. When we started, we only had the baseline that came with the simulator. As we experimented and learned, we documented ways to continue using the simulator beyond the competition. Post-NeurIPS, we had a robust codebase, knowledge, and new ideas, setting the stage for future projects.
The competition was the beginning, but later more than 5 research papers used the competition as a motivation and to perform experiments.
7. Skin in the game

“How much you truly ‘believe’ in something can be manifested only through what you are willing to risk for it.” Nassim Taleb
Before OpenReview, you could submit a paper, and if it did not get accepted, you got feedback. It could be terrible work, but still, you got feedback and no one needed to know about that terrible work you submitted. There was no penalty for submitting something incomplete. There was no skin in the game.
Competitions are different — by definition, there are winners, losers, and prizes. There is more at stake. Having skin in the game makes things more interesting and helps you learn in a distinct way.
Looking back, the competition gave me practice with decisions that kept appearing in industry: choosing an approximation, checking whether an improving training metric meant better performance, and deciding whether another experiment was worth the time.
The technical problems change, but I still find myself asking similar questions. What can we simplify? What are we leaving out? And what evidence would justify spending more time on it?
That is why I would recommend trying a competition. You get to build something, make those decisions with a team, and see how the result performs.