A mobile game can be technically ready for launch long before it is commercially ready.
That gap is where market testing matters.
Before committing significant development resources, UA budgets, or a global launch strategy, game teams can use paid acquisition to answer a more fundamental question: does the market respond to this game strongly enough to justify taking it further?
For experienced UA teams, this is not simply a question of finding a low CPI. A market test can reveal whether the game's concept is attractive, which creative propositions resonate, whether the store experience converts that interest, and (once a playable build exists) whether the users attracted by those messages actually behave like players worth acquiring.
There is also no universal answer to the question of where this test should happen. Some teams start with Meta. Others use TikTok, AppLovin, Google, or a combination of social and in-app networks. Some test the concept before the game is fully built; others wait until they have enough of the product to evaluate retention and monetization alongside acquisition.
The right choice depends on what the team is trying to learn.
That distinction is important because a marketing test is not necessarily a soft launch, and a soft launch is not necessarily the first marketing test. They answer different questions at different stages.
What a mobile game market test is actually trying to validate
The earliest marketing test is primarily a demand and marketability test.
At this stage, the team is asking questions such as:
Does the concept generate enough interest to earn a click?
Which game fantasy, mechanic, character, or value proposition attracts attention?
Does the creative communicate the concept clearly enough?
Is the interest strong enough to translate into an install or store action?
Are there meaningful differences between audiences, platforms, or markets?
Does the concept appear commercially promising relative to the cost of acquiring attention?
Once there is a playable or near-production build, the questions become more consequential:
Do users acquired by the strongest creative actually retain?
Does the game's first session deliver on the promise made by the ad?
Does the monetization model support the acquisition cost?
Are early CPI and engagement signals pointing toward viable LTV?
Does performance hold when the test moves beyond one audience or one source of supply?
This is why CPI alone is rarely enough to make a kill-or-continue decision.
CrazyLabs, for example, has described running highly granular marketability tests at the prototype stage, including CPI tests across Facebook and TikTok and, in some cases, an SDK-based network. While CPI was its primary initial signal, the team also examined Day 1 retention and 24-hour playtime so that a game with an apparently weak acquisition metric but unusually strong engagement would not automatically be discarded. (Pocket Gamer)
That is a useful distinction: the purpose of the test is not necessarily to predict the exact economics of a global launch. It is to reduce uncertainty around whether the game deserves the next investment.
Marketability testing and soft launch are different tests
The terms are sometimes used interchangeably, but they describe different levels of validation.
A marketability test can happen while a game is still a prototype or vertical slice. The product may not be ready for meaningful retention analysis. The objective is primarily to test the market's response to the idea and its presentation.
A soft launch happens with a live, playable product and real users. It can evaluate acquisition alongside retention, engagement, monetization, technical stability, and longer-term LTV signals.
This distinction is reflected in industry practice. Unity describes soft launch as a way to test retention and monetization while learning which creatives and supply channels produce higher-quality users. (Unity)
Adjust similarly describes soft launches as controlled releases designed to generate real-world data on engagement, monetization, technical performance, and UA efficiency before scaling. (Adjust)
So, for a new game, the testing process is better thought of as a progression:
Concept → marketability → playable build → soft launch → global launch
Not every game needs every stage in exactly this form. But separating the questions helps prevent teams from asking a prototype test to answer questions it cannot answer.

Start with the hypothesis, not the ad platform
One of the easiest mistakes in early testing is starting with:
"Should we test this on Meta or AppLovin?"
A better starting point is:
"What uncertainty are we trying to remove?"
If the question is whether the game concept itself has broad appeal, the test should be designed around creative and audience response.
If the question is whether a particular game fantasy is stronger than another, the test should isolate the propositions.
If the question is whether the game can acquire users economically in a particular market, then the test needs to introduce more realistic market and platform conditions.
If the question is whether users acquired cheaply are actually valuable, then the game needs to be sufficiently developed to measure downstream behavior.
This also changes how results should be interpreted.
A low CPI may indicate strong creative-market fit. It may also indicate that a particular network can find inexpensive users for that specific creative. A high CPI may indicate weak demand,but it could equally reflect an expensive audience, an unsuitable platform, weak creative execution, or insufficient learning.
The test design determines what can legitimately be concluded from the result.
Choosing where to run the initial test
There is no single universally superior testing channel. The major platforms expose games to different environments, audiences, optimization systems, and creative formats.
That means channel selection should be viewed as part of the experiment, rather than simply a media-buying decision.

Meta: a common starting point for concept and creative testing
Meta has historically been widely used for early mobile game testing because it provides a large audience and a mature ecosystem for testing creative propositions.
It can be particularly useful when the question is:
"Which version of this game idea makes people stop and respond?"
A team can test different hooks, game fantasies, characters, visual styles, or positioning while keeping other elements relatively consistent.
The important point is not that Meta is inherently the "best" market-testing platform. Rather, its scale and social-feed environment make it useful for testing the ability of a creative proposition to generate demand.
Meta itself highlights playable ads as a way to give prospective players an interactive preview before installation, describing the format as a way to attract higher-intent users. (Facebook)
That becomes particularly relevant for games where the mechanic itself is the proposition.
A playable can test something that a static ad or video cannot: will the user actually interact with the mechanic when given the opportunity?
But there is an important limitation. A result obtained on Meta should not automatically be treated as a universal measure of market demand. It is a measurement of demand within Meta's particular audience, inventory, creative environment, and optimization system.
TikTok: useful when creative-native demand is part of the question
TikTok introduces a different environment because creative behavior is more closely tied to entertainment and creator-led content.
For games whose proposition can be communicated through short-form video, humor, gameplay moments, characters, or creator-style storytelling, it can be valuable to test whether the concept works in a more entertainment-native context.
TikTok's own gaming research emphasizes a test-and-learn approach rather than a fixed campaign formula. Its examples include testing different optimization objectives, audiences, and creative approaches to identify what works for individual games. (TikTok For Business)
The broader point is that creative performance can be platform-specific.
A concept that produces strong CTR on a social feed does not necessarily produce equivalent performance inside gaming inventory. Conversely, an asset that looks mediocre in a social environment may work when the user is already in a game-playing mindset.
AppLovin and other in-app networks: testing in the environment where games are played
The argument for testing on a gaming-focused network is different.
AppLovin's own gaming proposition explicitly positions soft launch as a stage for testing marketability, identifying strong creatives, and finding audiences that convert. (AppLovin)
The value here is not simply another source of impressions. It is different supply and user context.
Sensor Tower's 2025 gaming advertising research found that AppLovin's share of mobile-game advertising grew substantially, while playables showed particularly strong impression growth on AppLovin and other in-app networks. (Sensor Tower)
That makes in-app inventory worth considering when the team wants to understand how the concept performs specifically in a gaming environment.
AppLovin also takes a different approach to creative testing. Its current guidance emphasizes creative volume and diversity, with the platform automatically testing combinations and using prediction before substantial delivery. (AppLovin)
For an experienced UA team, this creates an interesting testing consideration:
Do you want to test the creative, or do you want the platform to help discover which creative works?
Those are not exactly the same experiment.
Google: useful when the test needs broader automated inventory
Google App campaigns operate across Search, Google Play, YouTube, Discover, Display, and other Google inventory, with automated systems determining where and how assets are served. (Google Support)
This can make Google useful for validation, but it also changes the nature of an early experiment.
Because App campaigns automatically test combinations of assets and placements, it can be harder to interpret a simple "creative A versus creative B" comparison unless the experiment is deliberately structured.
Google now provides Directional Experiments for App campaigns, allowing advertisers to test groups of assets head-to-head against a selected success metric. (Google Support)
The broader implication is that teams should understand how much control a platform gives them over the experiment before using it to answer a narrow question.
One channel or multiple?
This is probably the most important strategic choice in the initial test.
Testing on a single channel has an obvious advantage: less noise.
If the objective is to determine which of five concepts has the strongest response, keeping the media environment constant can make the comparison cleaner.
But single-channel testing has a major weakness: platform bias.
A game can look highly promising because its creative happens to align with one platform's audience and distribution mechanics. That does not necessarily mean the game has equivalent demand across the broader market.
This is why some studios deliberately use multiple sources even at the early marketability stage.
CrazyLabs has described testing on both Facebook and TikTok and occasionally an SDK network to strengthen its marketability validation. (Pocket Gamer)
There is a trade-off:
One channel: cleaner experiment, lower complexity, faster learning.
Multiple channels: broader validation, more diverse user environments, but more variables to interpret.
Neither is automatically correct. The decision should follow the question being tested.
The creative test is often the real market test
For mobile games, the ad is not merely a delivery mechanism.
It is part of the product's market proposition.
The same game can be presented as:
a satisfying mechanic,
a progression fantasy,
a character-driven experience,
a social challenge,
a strategic problem,
a power fantasy,
a humorous scenario,
or a "fail and try again" challenge.
Those are effectively different products from the user's perspective.
That makes early creative testing particularly valuable.
AppsFlyer's 2025 State of Creative Optimization analyzed 1.1 million creative variations across 1,300 gaming and non-gaming apps. In gaming, the top 2% of creatives accounted for 53% of total ad spend, illustrating how strongly performance can concentrate around a small number of winners. (AppsFlyer)
At the same time, the industry is producing more creative than ever. AppsFlyer's 2026 State of Gaming for Marketers reported that top gaming advertisers were producing roughly 2,400–2,600 creative variations per quarter in 2025, up 25–30% year over year. (AppsFlyer)
The implication for an initial test is not simply "make more ads."
It is to make meaningfully different hypotheses.
Changing the font or CTA may tell you very little about market demand. Testing fundamentally different interpretations of the game can tell you much more.
For example:
Concept A: "Build the ultimate kingdom."
Concept B: "Can you survive the next 30 seconds?"
Concept C: "Solve the impossible puzzle."
The objective is not necessarily to identify a final production creative. It is to discover which promise creates the strongest response from the intended audience.
Look beyond CTR and CPI
Early metrics are useful precisely because they are early. But they also have limits.
A useful testing hierarchy is:
Attention
CTR, thumb-stop behavior, video engagement, IPM and related signals can indicate whether the proposition earns attention.
Intent
Click-to-install conversion, store conversion, playable engagement, or other post-click behavior can indicate whether the initial interest survives closer inspection.
Product response
Activation, tutorial completion, session length, early engagement, Day 1 retention and subsequent retention indicate whether the product delivers on the acquisition promise.
Economic response
Payer conversion, ARPU, ARPDAU, early ROAS, predicted LTV and payback begin to answer whether acquisition can ultimately work economically.
The further down the funnel the test goes, the more expensive and time-consuming it becomes,but also the more meaningful the signal.
This is why an experienced team should resist treating a single CPI threshold as a universal kill criterion.
PocketGamer's reporting on soft-launch practice highlights retention, ARPDAU, Day-x ARPU, ratings, CPI and CPE as complementary indicators rather than relying on a single KPI. (Pocket Gamer)
Similarly, Adjust's current soft-launch framework emphasizes activation, retention, engagement, CPI, cost per retained user, payer conversion and early ROAS. (Adjust)
The key is sequencing the metrics according to what the product is capable of proving at that stage.
Test the audience, not just the media
A common testing mistake is treating "cheap installs" as evidence of product-market fit.
They are not.
A low-cost audience that has little resemblance to the eventual target audience can produce misleading retention and monetization data.
This has been a recurring theme in soft-launch guidance. PocketGamer's analysis of soft-launch discipline argues that teams should begin with their intended audience rather than simply buying inexpensive random installs, because the resulting behavior may not represent the players the game ultimately needs to satisfy. (Pocket Gamer)
Market selection therefore matters.
A test market should be evaluated according to what it is supposed to represent:
purchasing power,
device mix,
genre penetration,
player behavior,
cultural relevance,
monetization environment,
advertising costs,
and similarity to the eventual priority markets.
Adjust similarly recommends selecting test markets that resemble the intended launch audience rather than choosing markets simply because they are inexpensive. (Adjust)
This is especially important for games with significant IAP economics. A market with cheap acquisition but dramatically different spending behavior can make early LTV estimates difficult to interpret.
The creative-to-product consistency check
One of the most valuable things an initial test can reveal is whether the marketing promise and the actual game experience match.
This becomes especially important for games using highly exaggerated advertising concepts.
A creative can generate excellent acquisition metrics by promising an experience that the actual game does not deliver. The result may look strong at the top of the funnel but collapse at activation or retention.
This is why the progression from marketability testing into soft launch matters.
The initial test asks:
"Do people want this?"
The product test asks:
"Do people who wanted this actually enjoy it?"
The strongest games have alignment between the two.
Google's open-testing approach explicitly positions pre-launch testing as a way to evaluate stability, retention, monetization and creative effectiveness with real users before launch. (Google Support)
That makes the transition from marketing test to live test particularly important: the more realistic the product becomes, the more downstream metrics should influence the decision.
What a strong initial test should leave you knowing
A good market test does not necessarily produce a definitive "launch" or "kill" answer.
It should instead reduce several important uncertainties.
By the end of an initial test, the team ideally has stronger evidence around:
1. Marketability
Does the concept generate meaningful demand?
2. Creative proposition
Which game fantasy, mechanic, visual identity, or message produces the strongest response?
3. Audience fit
Which audiences appear most responsive?
4. Platform behavior
Does the response appear consistent across different acquisition environments?
5. Conversion quality
Does attention translate into meaningful downstream action?
6. Product promise
When a playable build is available, does the game deliver what the ad promised?
7. Early economics
Do acquisition costs and early value signals point toward a potentially viable business model?
8. Next-stage uncertainty
What still cannot be answered without more users, more gameplay data, or longer cohorts?
That last question is often overlooked.
A test is useful not only when it produces a winner. It is useful when it tells the team what it still does not know.
There is no universal "best" market-testing channel
The current mobile advertising landscape makes a one-size-fits-all testing recommendation increasingly difficult to justify.
AppsFlyer's 2025 Performance Index analyzed 16.2 billion installs across 39,000 apps and 88 media sources. It found Google and Apple retaining leading positions while AppLovin, Meta, TikTok and other networks continued to strengthen their positions across gaming and non-gaming. (AppsFlyer)
Sensor Tower's 2025 gaming research similarly showed meaningful shifts in where mobile games were acquiring attention, including substantial growth in TikTok's and AppLovin's share of gaming advertising. (Sensor Tower)
That matters for testing because the channel landscape itself is changing.
A marketability test designed around one platform may answer a very specific question about that platform. A cross-platform test may provide broader validation, but at the cost of additional variables and budget.
For UA leaders, the useful question is therefore not:
"Which channel should we use to test our game?"
It is:
"Which environment,or combination of environments,will give us the most useful evidence for the decision we need to make next?"
That is a fundamentally different question.
The best initial test is the one that produces a decision
Market testing is not about accumulating dashboards.
It is about reducing uncertainty before the cost of being wrong increases.
For an early prototype, that may mean discovering that one game concept consistently attracts attention while another does not. For a playable build, it may mean discovering that a cheap acquisition source produces weak retention, while a more expensive source brings users who monetize better. For a soft launch, it may mean finding that the game's retention is strong but the monetization model cannot support the expected acquisition cost.
Each finding changes the next decision.
The most effective teams therefore treat market testing as an iterative evidence-gathering process, rather than a single campaign with a predetermined KPI threshold.
The platform matters. The market matters. The creative matters. The product matters. And the measurement framework matters.
But the most important variable is the question being asked.
A Meta-only test, an AppLovin test, a TikTok test, a Google App campaign, or a multi-network experiment can all be valid approaches when they are designed to answer the right question.
What matters is knowing exactly what the result can, and cannot, tell you before making the next bet.



