A Conversation with Macaron Founder Chen Kaijie: RL + Memory Makes the Agent a User's Exclusive "Doraemon"

 

RL infrastructure is difficult to standardize like cloud services.



oversea: Is RL currently the most crucial element for building agents? In your practical work, where do you see the most significant improvements and areas requiring deeper exploration?


Macaron Optimizes Memory with RL

Kaijie Chen: Most companies can't handle RL. We initially ran RL on a 70B model, but its memory capabilities for writing fiction were insufficient—it required reinforcement to produce 100,000 or 200,000 words. Later, our team continued exploring. This year, following the r1 paradigm—especially after Deepseek's 0528 release—we migrated RL from a 70B model to a 671B-scale model.


In China, very few teams can independently develop RL on 671B-scale models—probably fewer than five. Even many teams capable of pre-training, like Zhipu, cannot handle RL on dense 671B models. Most companies working on RL remain within the 10-200B range, where 200B is a watershed. Models under 200B can still be trained on a single GPU node; larger models necessitate expert parallelism and pipeline parallelism, distributing training and inference across multiple machines for simultaneous execution.


Consequently, most agent companies in the market claim RL isn't important—simply because they lack the capability to tackle it and struggle to even conceptualize the challenge. We've already observed diminishing returns in pre-training, signaling our entry into the “second half of intelligence enhancement.” RL is central to this phase. RL demands highly specific scenarios—there's no “super-general” RL domain that simultaneously boosts all capabilities of a universal agent. For instance, Anthropic has extensively applied RL in coding, resulting in significantly higher code quality compared to GPT-5. Other companies will optimize RL within distinct scenarios.


Leading companies must possess authentic models within selected scenarios, generate genuine content, obtain real user feedback, implement actual modifications, and iterate their RL outcomes. Thus, for RL development, models and scenarios must be intrinsically aligned. Creating a universal model today is exceptionally challenging. However, within specific domains—like personal agents—RL can enable Macaron to generate output surpassing all other models on the market. In such specialized scenarios, RL's value lies in pushing intelligence to its absolute limits within that context. Crucially, RL must operate on 700B-scale models. Attempts on 200B models currently struggle to cross the threshold toward AGI.


Therefore, with this market positioning, RL becomes critically important. Developers must continuously unlock the upper limits of intelligence.


Manus AI focuses on Context Engineering, believing this accelerates iteration speed—potentially releasing a new version every two to three days. I firmly believe products must iterate rapidly, growing alongside users. Yet RL can achieve this too.


Macaron proposes an approach called all-sync RL, compressing training time from “weekly” to “daily” cycles. A meaningful RL iteration can now complete in approximately 30 hours. This means our RL can also iterate on a daily basis. Applying this all-sync methodology allows us to both push the upper limits of intelligence and keep pace with product development cycles. For Macaron, RL is of paramount importance.



Macaron's all-sync mode achieves low GPU optimization consumption


oversea: Does the all-sync RL approach resemble how TikTok collects user feedback during the day, fine-tunes models overnight, and delivers more user-understanding models the next day? Can Macaron achieve this rhythm through RL?


Kaijie Chen: All-sync RL is an optimization algorithm for the training side, not the inference side. The training process has two parts: training and inference. During training, the model continuously infers and evolves, creating alternating GPU bubbles—essentially waste. Many optimization methods aim to eliminate this “bubble.”


We implemented “all-sync”: leveraging communication and model compression for better scheduling, enabling simultaneous training and inference with several-fold efficiency gains. Whereas 512 GPUs were previously needed for training, we can now complete initial training with just 48 GPUs.


Regarding online/on-policy, on-policy is achievable today; online is possible in the future. Both are crucial for user interaction systems.


• On-policy: The standard definition is that training data is entirely generated by the model itself. For example, Macaron's sub-agents are all written by the coding agent, ensuring clear optimization direction and more stable convergence.


• Online: Refers to real-time feedback loops between the model and user, enabling continuous online evolution.


On-Policy, All Async, All Sync RL Capabilities


oversea:RL enables small models to achieve what was previously only possible with large models. Do you think RL will still require large models in the next two to three years, or five years? Could it become “democratized,” allowing more AI developers to work with RL?


Kaijie Chen: RL infrastructure companies may emerge in the future, making RL accessible to all. However, several points warrant attention:


• Models shouldn't be too small. For specific small-scale scenarios, directly invoking large models may be more cost-effective than training a 20–70B model with RL; unless edge or local deployment is required, where RL combined with 10B, 20B, or 30B models might yield unique benefits;


• For non-edge large-scale scenarios, I believe models in the “600–1000 billion parameter range” better achieve AGI-level effects and offer greater scenario optimization.


The challenge with RL infrastructure lies in requiring real user feedback, retraining models, and finally deploying them. It's difficult to standardize like today's cloud services. Often, insufficient training data or poor data cleaning yields no results; It requires deep analysis + dozens of experiments to converge. We've never seen RL succeed on the first try. So how to price this service—multiple tiers or $2 million per project—and how to prove its value remain challenges. For Macaron, in-house development offers greater flexibility and scalability. 


oversea: Recently in Silicon Valley, I've observed numerous RL infrastructure companies—at least ten to twenty. Generally, they fall into two categories:


• More focused on data observation and services, helping agent developers label trajectories as thoroughly as possible;


• Directly building environments, such as providing coding agents with an environment that can generate CRMs/IDEs, offering greater automation and scalability.


What tasks do you think RL is particularly suited or unsuited for? What real opportunities does it bring to application developers?


Kaijie Chen: At Macaron, we're currently highly focused on two scenarios: memory and sub-agent writing, where training yields significant results. My personal experience shows: achieving 50–80% in a scenario is largely attainable through engineering and more complex context engineering; even 80–85% is manageable. But the leap from 85% to 95% is where RL delivers the most substantial improvement. At this point, continuing to pile on context engineering becomes like playing whack-a-mole: the process gets overly complex. Fixing one issue might introduce new problems elsewhere, making the system difficult to maintain. That's why we typically switch to RL in the final stages.


If the environment is well-defined, most tasks can be tackled with RL: drafting legal documents, conducting patient surveys, and so on. But if the environment isn't properly defined, and too much context engineering was piled on early on, you'll end up needing to supplement many technologies, requiring dismantling and end-to-end restructuring. Thus, the real-world environment is crucial.


oversea: As a DOTA player, when I see five DOTA NPCs charging at me, I feel they're fundamentally different from humans. Moreover, their strategies can outperform human players at the AI level. In scenarios with clear win/loss conditions, RL can already achieve 100% or higher—that's superhuman. Do you think there's potential for such “human-surpassing capabilities” to emerge in the future?


Kaijie Chen: Asking “where it outperforms humans” is really asking: how many people does my agent equate to in terms of thinking and executing this task?


• In the memory domain, Macaron will inevitably surpass humans. It will understand you better than your friends or even family—that's certain.


• Regarding coding agents, we'll likely measure its value in terms of human efficiency. It's difficult to say whether it can achieve limitless goals like “writing a WeChat, a TikTok, or even a ChatGPT.” It's not impossible, but that's beyond Macaron's current development objectives. Ultimately, it's more about: Can it create a Pomodoro timer as effective as the “million-rated Pomodoro timer” on the app store, but customized to your unique personality? Macaron could equate to: three front-end developers + three back-end developers, earning 30,000 yuan monthly (in Shenzhen), working together for two months. It remains something a small team can build—highly refined and deeply aligned with user needs. I believe this represents its ultimate achievable state.


oversea: I constantly forget friends' birthdays—Macaron would definitely outperform most friends like me. Meanwhile, average users have limited access to IT outsourcing for task solutions. Here, Macaron delivers far better results than typical user-accessible solutions.


Kaijie Chen: Exactly. Moreover, Macaron's Coding Agent and Anthropic's Coding Agent require very different optimization approaches. Anthropic focuses on complex, large-scale programming scenarios—it's an engineering solution. Macaron prioritizes small, complete functionalities and ensures they run correctly on the first try. Macaron can make mistakes, but what it generates must be usable. Users should see the finished project Macaron created for them the very first time they open it.




Small, personal apps meet real needs


oversea: When you launched Macaron, you posted on your social media that you were previously a Duolingo user but canceled your Duolingo subscription after Macaron's release, believing it could already serve as an excellent language learning assistant. So, I'm curious: as Macaron undergoes reinforcement learning with the community and users, do you plan to deepen its capabilities in specific verticals or focus first on building a more general-purpose Personal Agent? How will you balance these approaches?


Kaijie Chen: Over a week after launch, over 7,000 users created more than 10,000 diverse mini-apps. Among these, certain major categories began to emerge. This relates to our initial strategy. In the first two weeks, users often felt: “Macaron promises to do everything, but nothing feels truly polished.” This presented me with a tough trade-off:


  1. Only promise what I know I can do well, limiting the scope to what I can execute effectively;
  2. Or keep the door wide open. The downside is subpar execution, but the upside is gathering genuine user intent distribution.

I chose the latter to uncover authentic needs. Next, we can identify what works well and what doesn't, drawing an intersection between user demand and Coding Agent's capabilities as the first step for enhancement. Among the over ten thousand mini-apps, a clear convergence is emerging in Tracker/Planner (life logging and planning), such as:


• Diet and fitness tracking;


• Mood journals, game logs, academic progress tracking, portfolio organization;


• Some users log “the color and shape of their poop”;


• Many Japanese users enjoy generating piano and guitar sheet music (Possibly tied to local music market pricing).


Macaron will feed these tracking/planning use cases into the next RL training phase to refine them. The chatbot will moderately incorporate these needs but won't rigidly constrain itself, as we need to observe where genuine user demands truly land next. AI evolves rapidly—I believe we can progressively unlock more needs and push boundaries outward.


I believe focusing on a specific vertical is also a dynamic exploration journey with users, somewhat like Xiaohongshu initially focusing on overseas shopping before expanding into beauty and camping. For Macaron, this will happen faster since production is entirely in-house—no need to coordinate an external ecosystem. We just need to enhance the Agent capabilities.


Macaron Feature Showcase


oversea: The community feedback on Macaron is fascinating—both positive and controversial. I've highlighted two key points:


  1. Users feel the cloud-optimized mini-programs are too slow or unusable;


  1. Users perceive Macaron as having a “motherly” vibe, like it constantly wants to help.


Do these reactions align with your expectations? How will you address them moving forward?


Kaijie Chen: This feedback aligns perfectly with our expectations. Macaron's initial strategy was to broadly collect user demand distributions while identifying model failure cases. During beta testing, we knew we couldn't cover every scenario ourselves—we were prepared for criticism. And we got it. But I completely understand.


The tone of user comments largely depends on the first mini-app they try: If their first attempt lands squarely within the model's capabilities, they're often amazed, even moved—some users have even cried. But if that first app fails, like when links don't work, they lose faith in vibe coding, thinking AI just can't deliver today.


So I'm incredibly grateful to our early adopters for trusting us and giving it a shot. The solution is actually quite clear:


• First, identify the intersection of supply and demand, narrowing the scope to what we can execute well;


• We launched an update on August 29th to improve speed and stability;


• The most urgent user requests, like tracker/planner (recording and planning) use cases, will be prioritized in our next RL training phase.


I believe the model's unlocking pace is rapid—in six months or a year, it will be vastly different from today. We're also testing new generative capabilities that will only grow stronger.


Regarding “motherly warmth,” Macaron still needs training on key conversational points. While we grasped narrative dialogue in previous virtual characters (like MidReal), there's room for improvement in “Doraemon-style” daily companionship chats. You'll notice some differences in versions released after August 29th.


oversea: Have any user interactions surprised you with particularly compelling use cases?


Kaijie Chen: Several instances have genuinely amazed me.


• Golf swing analysis: Yesterday, a friend wanting to learn golf created a practice log. He requested not just static records but “motion video analysis” to improve accuracy. I thought this case would definitely fail, but Macaron actually delivered: after the user uploaded a video, it analyzed the swing and provided improvement suggestions. This wasn't a scenario we specifically trained it for, yet it extrapolated from existing knowledge—which I find incredibly cool.


• Moving Planning: A user in her 40s recently relocated to a new city in the US. She used Macaron to plan her new life: where to explore, how to schedule activities, and how to practice mood-boosting yoga and meditation. Seeing this made me genuinely happy. People navigating life changes—like starting school, moving, getting a pet, or having a baby—really need a helper.


• Family Recipe Management: One user built a family recipe app for grandparents and children with different tastes. He designed a Dianping-like feature where family members rate dishes. The system automatically filters the most suitable recipes based on who's home that day, organizing ingredients and instructions. This deeply personal need is both practical and fascinating.


These use cases rarely find ready-made solutions in app stores because they're too personalized. Many ask: “Does an app need to be ten times better than what's in the App Store to attract users to Macaron?” My answer is: Absolutely not. Sometimes it's just about building a tool for yourself or a small group—one that perfectly fits your specific needs.


Macaron User Feedback


oversea: Applications like this used to be developed only when the market was large enough and had the greatest common denominator. It's like the vision Sam Altman or Dario Amodei often talk about: with model intelligence, applications will flow like tap water—turn on the faucet and it pours out. In daily life, Macaron is that faucet, covering the last-mile scenarios.


Kaijie Chen: Exactly. Writing code now truly feels like turning on a tap—the cost is comparable. This empowers everyone to create things that align more closely with their vision. It's different from the past, where others built for you. Now, even ordinary users can craft tools that suit them better than anything else. The future holds immense potential.


Compared to work scenarios, layering life scenarios delivers greater value


oversea: Back to the business model. Long-term, do you envision building a new-era App Store? Or something fundamentally different?


Kaijie Chen: Honestly, Macaron isn't an App Store either. First and foremost, it's the user's Personal Agent. The commercial potential unlocked by Personal Agents is already substantial.


The difference between life and work is this: Overlaying life scenarios creates greater value; but in work scenarios—like one project over another—arbitrarily overlaying contexts can lead to disaster. When Macaron knows what sports you've been doing or what foods you enjoy, it tailors your fitness plan differently; when you travel, it recommends different destinations and hotels; gifts for friends or home decor for moving become personalized too. This layering of scenarios creates greater impact than pursuing perfection in isolated features.



Commercially, you can book flights, shop, order takeout; or use it as a fitness coach or nutritionist. This potential is already substantial.


I don't want Macaron to become a platform where developers agonize over what subagent others might need, then painstakingly customize and refine it before releasing it to everyone. Instead, I envision users sharing their authentic lifestyles—people who are simply themselves, sharing their stories naturally. This aligns with my vision for Macaron's community. A healthy community ecosystem and mindset are crucial for any product. What the founder champions ultimately determines the product's market position and longevity. So I aim to preserve this ethos—one of “lifestyle sharing,” not “creation and consumption.”


oversea: Many Agents now lean toward subscription-based business models, essentially functioning like Office—paying for “productivity” over a set period. If Macaron ultimately aims to build a community, wouldn't that rule out this path?


Kaijie Chen: Macaron currently operates on a subscription model. Subscriptions aren't inherently undesirable: it's reasonable to pay a monthly salary for a “personal assistant.” In the future community, when your Agent is shared and used by more people, you should also receive some return. Additionally, the Personal Agent section could incorporate ads or third-party integrations. There's actually significant room for business model innovation here.


oversea: How was the decision made to transition from MidReal to building Macaron? How long did it take?


Kaijie Chen: On one hand, user observation: users immersed in MidReal's virtual world grew increasingly conflicted. We wanted to positively impact real life. On the other hand, the story direction seemed misaligned with today's optimal path for model capability upgrades—models became better at writing code but worse at crafting narratives (especially noticeable with GPT-5). Therefore, pursuing a path not aligned with the fastest advancements in intelligence might run counter to the current trajectory of AI development.


Regarding serving life, the first venture I pursued after taking a leave of absence from Duke in 2018 was a home smart robot—a project I started with my current cofounder. We both share a vision for enriching daily life and pursue interesting lifestyles ourselves, which motivated us to take a step toward serving life. Simultaneously, we witnessed the emergence of models like Claude code, which can achieve genuinely superior coding capabilities, and Deepseek R1, capable of reinforcement learning training while also possessing strong coding abilities. These factors combined propelled us to transition from MidReal to a new vision.


oversea: From MidReal to the current Macaron, as the founder, what have been your biggest shifts in mindset and work approach?


Kaijie Chen: Previously, I believed securing a top-three position in a niche market—like MidReal did—could establish a solid market presence or a viable business. That was indeed my mindset. But the more I worked, the more I realized achieving top-three status in a small niche holds limited significance. I now aspire to pioneer truly meaningful innovations at the forefront of the AI wave. So my mindset has shifted: it's not about being top in a niche market, but entering a larger arena where everyone has a chance.


The Personal Agent space is unique. Many players will enter it, and it might eventually resemble the social media landscape: overseas, Facebook dominates, but Instagram thrives alongside Telegram, X, BeReal, and Snapchat—each platform with its distinct personality.


Personal Agents will follow a similar pattern: they'll be like friends with different personalities. You could have a highly professional friend, or a warm, almost “motherly” friend. That's the model I want to pivot towards. While Macaron emphasizes lifestyle, our work rhythm has actually entered a “lifestyle without life”—recently we've been focused on fixing user-reported issues and rapid iteration.


oversea: That sounds a lot like reinforcement learning: when rewards plateau in one area, you seek steeper reward curves elsewhere.


Kaijie Chen: Exactly. Our primary pursuit isn't wealth—it's more of a lifestyle choice. After hitting a ceiling in our original direction, we chose to pivot to this new path.


oversea: If we compare Personal Agent to Facebook/Instagram/Telegram, is ChatGPT currently in Facebook's position?


Kaijie Chen: GPT is undoubtedly in Facebook's position. It already has 400 million DAU today—it's unstoppable growth and the biggest open card. Many companies that weren't traditionally “risk-takers” are now entering the OpenAI space. It's truly powerful.


oversea:So with Macaron AI, is your timing too early or too late? Or is now the perfect moment?


Kaijie Chen:The timing is spot on.


• Micro-level: After ChatGPT 5 launched, Macaron faced significant criticism (which actually worked in our favor);


• Macro-level: We're among the first teams to launch a Personal Agent, allowing us to capture user awareness early. I still have at least three to six months of breathing room.


I find it hard to imagine ChatGPT quickly pivoting its product to do this. While GPT is powerful and can handle some personal needs, it won't approach Personal Agents the way we do. That direction remains quite distant for them. If they were to enter this space, they'd need significant product adjustments—both in interaction design and user perception.


I actually don't foresee intense competition. I'm confident that once Macaron builds its own community, both Macaron and OpenAI will have distinct user bases. Users might even have both apps on their phones, using each in different contexts. If someone opens ChatGPT every time they work, I doubt they'd feel comfortable chatting with GPT like a friend at home afterward—it feels a bit disjointed or like a split personality. Therefore, Macaron can coexist with GPT and even compete against it.


oversea: Standing alongside giants takes real courage. Currently, ChatGPT has largely occupied the “colleague” position in users' minds. Once that space is secured, a new mental space will inevitably open up for all startups, including Macaron. The agent domain is vast—some focus on general-purpose solutions, others on verticals, and still others on infrastructure. Which directions do you think are underestimated, and which are overestimated?


Kaijie Chen: Currently, many small agents building workflows based on large models aren't sustainable and will eventually be overtaken by larger agents. For example, a Personal Agent will inevitably excel at “travel planning” because it has your lifestyle data—likely outperforming standalone travel agents. However, agents still hold significant opportunities in specialized scenarios.


Yes, writing code is now as accessible as running water, with comparable costs. This empowers everyone to create solutions that better align with their vision. It's fundamentally different from the past, where others built for you; now even ordinary users can craft tools that suit their needs better than off-the-shelf options. The future holds immense potential.

Comments

Popular posts from this blog

Alibaba's Z-Image-Turbo: How a 6.15B Parameter AI Model Crushed 20B Giants in Image Generation (0.8s & Perfect Chinese Text)

The Group Chat Was Already Dead by Message 12

Why AI Agents Need More Than Reusable Skills