[{"data":1,"prerenderedAt":377},["ShallowReactive",2],{"insights-list-en":3},[4,227,299],{"id":5,"title":6,"affiliation":7,"audiences":8,"author":12,"body":13,"date":212,"description":213,"draft":214,"extension":215,"featured":214,"meta":216,"navigation":217,"path":218,"seo":219,"stem":220,"tags":221,"updatedAt":212,"urlname":225,"__hash__":226},"insights_en/en/insights/dyna：讓機器人用「預測影片」學會動作的模型.md","Dyna: A Model That Teaches Robots to Act via \"Video Prediction\"","Department of Computer Science, National Tsing Hua University",[9,10,11],"application","researcher","developer","Wan-Chi Chang",{"type":14,"value":15,"toc":208},"minimark",[16,24,30,33,70,73,87,93,98,111,116,144,149,191,196],[17,18,19,23],"p",{},[20,21,22],"strong",{},"What is Dyna?"," Dyna is a leading representative of the World-Action Model (WAM) approach. Unlike VLA (Vision-Language-Action) models, which \"understand language and then generate actions,\" Dyna uses a video diffusion model to simultaneously predict \"how the next frame will change\" and \"how the robot should move in the next step.\" Action generation is merely a shallow branch connected next to the video backbone.",[17,25,26,29],{},[20,27,28],{},"Why Train Alongside Videos?"," Dyna's logic is that what robots lack is not linguistic common sense, but rather an intuition for how the physical world evolves—what happens when a hand touches an object, or how a rope gets knotted. This kind of knowledge cannot be learned from language models, but it is abundant in human first-person (egocentric) videos. Furthermore, such video data is practically unlimited and vastly cheaper than collecting demonstrations via teleoperated robotic arms. The researchers extract hand poses from videos to use as pseudo-action labels for training.",[17,31,32],{},"Experimental findings show that models trained solely on actions without generating video struggle to transfer to unseen robotic platforms. However, once \"simultaneous future video prediction\" is added, zero-shot performance significantly outperforms pure action training. Even simply feeding unlabelled human videos—used purely to learn video generation—continuously improves robot task performance. The videos themselves become a new source of data augmentation.",[17,34,35,38,39,49,50,57,58,65,66,69],{},[20,36,37],{},"Recent Developments"," In April 2025, ",[40,41,45],"a",{"href":42,"rel":43},"https://arxiv.org/abs/2504.02792",[44],"nofollow",[46,47,48],"em",{},"Unified World Models",", and in early 2026, ",[40,51,54],{"href":52,"rel":53},"https://arxiv.org/abs/2602.15922",[44],[46,55,56],{},"DreamZero",", integrated world models with action synthesis into a single architecture. In August 2026, Dyna Robotics released ",[40,59,62],{"href":60,"rel":61},"https://www.dyna.co/dyna-2",[44],[20,63,64],{},"Dyna-2",", pushing the pre-training scale to millions of hours of human egocentric video. This work validated for the first time that scaling human data yields power-law improvements and revealed a ",[20,67,68],{},"\"Human-to-Robot Transfer Scaling Law\"",": even without seeing any robot data, expanding human video data alone leads to a simultaneous drop in offline prediction error for robot tasks. On physical robots, with just a few hours of robot demonstrations for post-training, the million-hour pre-trained model achieved top performance across dual-arm, dexterous hand, and semi-humanoid platforms. It even learned to open bottle caps with only ten minutes of teleoperated data, demonstrating high robustness against lighting changes, visual occlusions, and persistent disturbances.",[17,71,72],{},"The Dyna approach is currently explored by only a few teams due to its higher training costs, and whether it can continuously scale to tens of millions of hours remains unknown. Nevertheless, it offers a distinct alternative to VLA: rather than borrowing linguistic common sense, it allows the model to directly learn how the world changes.",[17,74,75,76,79,80,83,84],{},"Most intriguing is a counterintuitive detail: feeding more pure video data consistently helps robot tasks, but has almost no effect—or even slightly degrades performance—when predicting human actions themselves. The same batch of data works for \"crossing over to another body,\" but is ineffective for \"staying in the original body.\" This suggests that video prediction learns not \"how to move,\" but rather \"how the world changes.\" Although human hands and robot grippers look completely different, physical rules—such as a pushed cup rolling or a pulled rope tightening—apply universally. If this path proves viable, the data bottleneck for robotics shifts from ",[46,77,78],{},"\"how many people are willing to teleoperate arms\""," to ",[46,81,82],{},"\"how many people's lives are recorded.\""," And that brings up a much harder question: ",[46,85,86],{},"whose lives, and with whose consent?",[17,88,89,90],{},"🔗 ",[20,91,92],{},"Further Reading",[17,94,95],{},[20,96,97],{},"Featured Work",[99,100,101],"ul",{},[102,103,104,110],"li",{},[40,105,107],{"href":60,"rel":106},[44],[20,108,109],{},"Dyna-2 Technical Report"," | Dyna Robotics, 2026",[17,112,113],{},[20,114,115],{},"World Models × Action",[99,117,118,126,134],{},[102,119,120,125],{},[40,121,123],{"href":42,"rel":122},[44],[20,124,48],{}," | Coupled Pre-training of Video and Action Diffusion",[102,127,128,133],{},[40,129,131],{"href":52,"rel":130},[44],[20,132,56],{}," | World Action Models are Zero-shot Policies",[102,135,136,143],{},[40,137,140],{"href":138,"rel":139},"https://arxiv.org/abs/2602.16710",[44],[20,141,142],{},"EgoScale"," | Scaling Dexterous Manipulation with Egocentric Human Videos",[17,145,146],{},[20,147,148],{},"Comparison: The VLA Approach",[99,150,151,161,171,181],{},[102,152,153,160],{},[40,154,157],{"href":155,"rel":156},"https://arxiv.org/abs/2307.15818",[44],[20,158,159],{},"RT-2"," | The Pioneer of the VLA Route",[102,162,163,170],{},[40,164,167],{"href":165,"rel":166},"https://arxiv.org/abs/2406.09246",[44],[20,168,169],{},"OpenVLA"," | Open-source Implementation, Ideal for Beginners",[102,172,173,180],{},[40,174,177],{"href":175,"rel":176},"https://arxiv.org/abs/2410.24164",[44],[20,178,179],{},"π0"," | Continuous Action Generation via Flow Matching",[102,182,183,190],{},[40,184,187],{"href":185,"rel":186},"https://arxiv.org/abs/2503.14734",[44],[20,188,189],{},"GR00T N1"," | NVIDIA's Humanoid Robot Foundation Model",[17,192,193],{},[20,194,195],{},"Technical Foundations",[99,197,198],{},[102,199,200,207],{},[40,201,204],{"href":202,"rel":203},"https://arxiv.org/abs/2210.02747",[44],[20,205,206],{},"Flow Matching"," | The Generative Method Currently Used by Both Approaches",{"title":209,"searchDepth":210,"depth":210,"links":211},"",2,[],"2026-09-14","Dyna represents a World-Action Model approach that leverages video diffusion models to simultaneously predict future frame evolutions and robot actions from vast human egocentric videos, learning physical world dynamics to overcome traditional teleoperation data bottlenecks and achieve highly robust, cross-platform robot manipulation with minimal fine-tuning.",false,"md",{},true,"/en/insights/dyna",{"title":6,"description":213},"en/insights/dyna：讓機器人用「預測影片」學會動作的模型",[222,223,224],"dyna","world-action model","video prediction","dyna-world-action-model-video-prediction","g9MxaZCZblyWg2B5vdKj2Po4U2vkJjH5Du4BMo-MHEc",{"id":228,"title":229,"affiliation":230,"audiences":231,"author":232,"body":233,"date":289,"description":290,"draft":214,"extension":215,"featured":214,"meta":291,"navigation":217,"path":292,"seo":293,"stem":294,"tags":295,"updatedAt":289,"urlname":297,"__hash__":298},"insights_en/en/insights/agent-skills：讓-ai-學會可重複使用的專業技能.md","Agent Skills: Teaching AI Reusable Professional Capabilities","Department of Statistics, National Taipei University (NTPU)",[9,10,11],"Wei-Chen Huang",{"type":14,"value":234,"toc":287},[235,246,257,263],[17,236,237,240,241,245],{},[20,238,239],{},"Agent Skills"," One of the most notable recent developments in Agentic AI is",[40,242,239],{"href":243,"rel":244},"https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills",[44],". In the past, using generative AI often required re-entering prompts, background information, and operational requirements every single time. Agent Skills, however, package workflows, rules, reference materials, and tools into reusable \"skills.\" When the AI encounters a suitable task, it can automatically load and execute them. This allows AI to go beyond answering questions to completing complex, repetitive tasks based on established workflows, enabling individuals and organizations to gradually build up their own AI capabilities. For enterprises, deploying the same set of skills ensures consistent workflows across teams, reducing training overhead and operational discrepancies.",[17,247,248,251,252,256],{},[20,249,250],{},"Progressive Disclosure"," A core concept behind Agent Skills is",[40,253,250],{"href":254,"rel":255},"https://blog.aihao.tw/2026/05/20/llm-knowledge-base/",[44],". The AI does not need to read the full content of all available skills upfront; instead, it maintains awareness of available skills and determines which one is needed based on the task at hand before loading relevant instructions, documents, examples, and resources. This approach minimizes context and token consumption while preventing information overload that could impact model performance. As the number of skills grows, this \"load-on-demand\" architecture makes it far easier to expand the agent's capabilities across different domains.",[17,258,259,262],{},[20,260,261],{},"Accessibility for Everyone"," Crucially, Agent Skills are not just for developers. Everyday users can turn routine, repetitive tasks into Skills—such as formatting meeting minutes according to company templates, organizing research papers based on specific rules, refining English resumes, generating weekly reports, or even structuring travel itineraries and self-study plans. Users mainly need to articulate the step-by-step procedures, rules, examples, and reference materials clearly without needing complex programming skills. Alternatively, they can leverage pre-built skills created by others and customize them as needed. In the future, proficiency in AI may shift from simply \"knowing how to write prompts\" to knowing how to structure expertise and workflows into reusable Skills, turning AI into a continuously evolving work assistant.",[17,264,265,268,269,274,275,280,281,286],{},[20,266,267],{},"Recommended Agent Skills"," To experience Agent Skills firsthand, you can start with popular, versatile projects on GitHub. For instance, the official",[40,270,273],{"href":271,"rel":272},"https://github.com/anthropics/skills?utm_source=chatgpt.com",[44],"Anthropic Skills","repository (with ~171k stars) includes skills for handling documents like PDFs, PowerPoint presentations, Word files, and Excel spreadsheets—ideal for daily office work and data organization, as well as serving as a direct reference for skill design. Another example is",[40,276,279],{"href":277,"rel":278},"https://github.com/mvanhorn/last30days-skill?utm_source=chatgpt.com",[44],"Last30Days","(~59k stars), which automatically searches Reddit, X, YouTube, TikTok, Hacker News, and web sources to summarize recent trends and discussions on a given topic, making it well-suited for market research and news tracking. Meanwhile,",[40,282,285],{"href":283,"rel":284},"https://github.com/blader/humanizer?utm_source=chatgpt.com",[44],"Humanizer","(~38k stars) focuses on refining text to remove common AI writing artifacts and achieve a more natural tone, proving highly useful for reports, articles, and general drafting. These cases demonstrate that Agent Skills are moving beyond developer toolkits and into research, writing, and everyday office workflows.",{"title":209,"searchDepth":210,"depth":210,"links":288},[],"2026-08-31","Agent Skills encapsulate workflows, rules, and domain context into reusable AI capabilities, enabling models to perform complex tasks through progressive disclosure without cluttering context windows.",{},"/en/insights/agent-skills-ai",{"title":229,"description":290},"en/insights/agent-skills：讓-ai-學會可重複使用的專業技能",[239,250,296,279,285],"Anthropic","agent-skills-progressive-disclosure","uQZYIvRAoFQ_XYgwcyfSJQGeEV2wvV_ZZ0ERPsXKaOU",{"id":300,"title":301,"affiliation":7,"audiences":302,"author":303,"body":304,"date":364,"description":365,"draft":214,"extension":215,"featured":214,"meta":366,"navigation":217,"path":367,"seo":368,"stem":369,"tags":370,"updatedAt":374,"urlname":375,"__hash__":376},"insights_en/en/insights/vla-robot-language-models.md","VLA: Teaching Robots to Understand Plain Language",[9,11,10],"Da-An Li",{"type":14,"value":305,"toc":359},[306,311,314,317,321,324,327,330,334,337,356],[307,308,310],"h2",{"id":309},"what-is-vla","What is VLA",[17,312,313],{},"VLA stands for Vision-Language-Action. It takes a camera image plus a natural-language instruction — for example, \"put the ketchup in the basket\" — and outputs what the robot should do next: joint angles for the arm, and whether the gripper opens or closes.",[17,315,316],{},"Traditionally, each robot task meant an engineer hand-writing a program, or training a task-specific model that had to be redone whenever the setting changed. VLA aims for \"one model, many tasks, many robots\" — a general-purpose brain for robotics.",[307,318,320],{"id":319},"why-build-on-a-vlm-backbone","Why build on a VLM backbone",[17,322,323],{},"It comes down to data. The web offers trillions of words and billions of images, but robot manipulation data has to be recorded one episode at a time by a human teleoperating a mechanical arm — several orders of magnitude less. Trained from scratch, a model simply memorises the demonstrations it was shown, and fails the moment it meets an unfamiliar cup.",[17,325,326],{},"Vision-language models (VLMs), by contrast, have already learned from web-scale data what ketchup looks like, that the red ball is on the left, and that a mug is grasped by its handle — visual common sense and language understanding. VLA takes that brain as-is and attaches an \"action expert\" behind it, treating actions as just another language to be generated.",[17,328,329],{},"The result is that a robot's ability to generalise is inherited from web data rather than squeezed out of scarce robot data. Only then does it stand a chance with objects it has never seen and phrasings it has never heard.",[307,331,333],{"id":332},"recent-developments","Recent developments",[17,335,336],{},"RT-2 first showed the approach was viable in 2023, and the open-source release of OpenVLA in 2024 put it within reach of academic labs. From 2025 onward the work turned toward products: Physical Intelligence's π0 / π0.5 use flow matching to generate continuous, high-frequency motion, while NVIDIA GR00T and Google Gemini Robotics adopt a dual-system design — a slow VLM for understanding and planning, and a fast action expert for real-time control.",[17,338,339,340,343,344,347,348,351,352,355],{},"Several threads define 2026. ",[20,341,342],{},"Efficiency",": inference on a seven-billion-parameter model is still an order of magnitude away from real-time control, driving a wave of distillation, quantisation, and token-pruning work. ",[20,345,346],{},"Reasoning",": letting the model \"think a step\" before it acts. ",[20,349,350],{},"Data scaling",": pre-training on tens of thousands of hours of first-person human video, with early evidence that dexterity follows a scaling law of its own. And ",[20,353,354],{},"world models",": having the model predict what happens next while it decides how to move.",[17,357,358],{},"VLA has not solved everything — generalisation remains brittle, and evaluation on real hardware is still expensive. But the direction is clear: take the recipe that worked for language models, and carry it into the physical world.",{"title":209,"searchDepth":210,"depth":210,"links":360},[361,362,363],{"id":309,"depth":210,"text":310},{"id":319,"depth":210,"text":320},{"id":332,"depth":210,"text":333},"2026-08-19","Bringing the recipe behind language models into the physical world — how VLA lets robots handle objects and instructions they have never seen before.",{},"/en/insights/vla-robot-language-models",{"title":301,"description":365},"en/insights/vla-robot-language-models",[371,372,373],"VLA","VLM","Robotics",null,"vla-robot-language-models","oQgxi1MJg2gSTrWMC2LgG08FI2MKMTZg8cCtdUlb-qI",1790125405182]