martes, 7 de diciembre de 2021

A system for designing and training intelligent soft robots

Let’s say you wanted to build the world’s best stair-climbing robot. You’d need to optimize for both the brain and the body, perhaps by giving the bot some high-tech legs and feet, coupled with a powerful algorithm to enable the climb. 

Although design of the physical body and its brain, the “control,” are key ingredients to letting the robot move, existing benchmark environments favor only the latter. Co-optimizing for both elements is hard — it takes a lot of time to train various robot simulations to do different things, even without the design element. 

Scientists from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), aimed to fill the gap by designing “Evolution Gym,” a large-scale testing system for co-optimizing the design and control of soft robots, taking inspiration from nature and evolutionary processes. 

The robots in the simulator look a little bit like squishy, moveable Tetris pieces made up of soft, rigid, and actuator “cells” on a grid, put to the tasks of walking, climbing, manipulating objects, shape-shifting, and navigating dense terrain. To test the robot’s aptitude, the team developed their own co-design algorithms by combining standard methods for design optimization and deep reinforcement learning (RL) techniques. 

The co-design algorithm functions somewhat like a power couple, where the design optimization methods evolve the robot’s bodies and the RL algorithms optimize a controller (a computer system that connects to the robot to control the movements) for a proposed design. The design optimization asks “how well does the design perform?” and the control optimization responds with a score, which could look like a five for “walking.” 

The result looks like a little robot Olympics. In addition to standard tasks like walking and jumping, the researchers also included some unique tasks, like climbing, flipping, balancing, and stair-climbing. 

In over 30 different environments, the bots performed amply on simple tasks, like walking or carrying an item, but in more difficult environments, like catching and lifting, they fell short, showing the limitations of current co-design algorithms. For instance, sometimes the optimized robots exhibited what the team calls “frustratingly” obvious nonoptimal behavior on many tasks. For example, the “catcher” robot would often dive forward to catch a falling block that was falling behind it.

Even though the robot designs evolved autonomously from scratch and without prior knowledge by the co-design algorithms, in a step toward more evolutionary processes, they often grew to resemble existing natural creatures while outperforming hand-designed robots.  

“With Evolution Gym we’re aiming to push the boundaries of algorithms for machine learning and artificial intelligence,” says MIT undergraduate Jagdeep Bhatia, a lead researcher on the project. “By creating a large-scale benchmark that focuses on speed and simplicity, we not only create a common language for exchanging ideas and results within the reinforcement learning and co-design space, but also enable researchers without stat-of-the-art computing resources to contribute to algorithmic development in these areas. We hope that our work brings us one step closer to a future with robots as intelligent as you or I.” 

In certain cases, for robots to learn just like humans, trial and error can lead to the best performance of understanding a task, which is the thought behind reinforcement learning. Here, the robots learned how to complete a task like pushing a block by getting some information that will assist it, like “seeing” where the block is, and what the nearby terrain is like. Then, a robot gets some measurement of how well it’s doing (the “reward”). The more the robot pushes the block, the higher the reward. The robot had to simultaneously balance exploration (maybe asking itself “can I increase my reward by jumping?”) and exploitation (further exploring behaviors that increase the reward). 

The different combinations of “cells” the algorithms came up with for different designs were highly effective: one evolved to resemble a galloping horse with leg-like structures, mimicking what’s found in nature. The climber robot evolved two arms and two leg-like structures (kind of like a monkey) to help it climb. The lifter robot resembled a two-fingered gripper. 

One avenue for future research is so-called “morphological development,” where a robot incrementally becomes more intelligent as it gains experience solving more complex tasks. For example, you’d start by optimizing a simple robot for walking, then take the same design, optimize it for carrying, and then climbing stairs. Over time, the robot's body and brain “morph” into something that can solve more challenging tasks compared to robots directly trained on the same tasks from the start. 

"Evolution Gym is part of a growing awareness in the AI community that the body and brain are equal partners in supporting intelligent behavior,” says University of Vermont robotics professor Josh Bongard. “There is so much to do in figuring out what forms this partnership can take. Gym is likely to be an important tool in working through these kinds of questions.”

Evolution Gym is open source and free to use. This is by design as the researchers hope that their work inspires new and improved algorithms in codesign. 

The work was supported by the Defense Advanced Research Projects Agency. Bhatia wrote the paper alongside MIT undergraduate Holly Jackson, MIT CSAIL PhD student Yunsheng Tian, and Jie Xu, as well as MIT Professor Wojciech Matusik. They are presenting the research at the 2021 Conference on Neural Information Processing Systems.



de MIT News https://ift.tt/3dyqNiY

lunes, 6 de diciembre de 2021

Technique enables real-time rendering of scenes in 3D

Humans are pretty good at looking at a single two-dimensional image and understanding the full three-dimensional scene that it captures. Artificial intelligence agents are not.

Yet a machine that needs to interact with objects in the world — like a robot designed to harvest crops or assist with surgery — must be able to infer properties about a 3D scene from observations of the 2D images it’s trained on.     

While scientists have had success using neural networks to infer representations of 3D scenes from images, these machine learning methods aren’t fast enough to make them feasible for many real-world applications.

A new technique demonstrated by researchers at MIT and elsewhere is able to represent 3D scenes from images about 15,000 times faster than some existing models.

The method represents a scene as a 360-degree light field, which is a function that describes all the light rays in a 3D space, flowing through every point and in every direction. The light field is encoded into a neural network, which enables faster rendering of the underlying 3D scene from an image.

The light-field networks (LFNs) the researchers developed can reconstruct a light field after only a single observation of an image, and they are able to render 3D scenes at real-time frame rates.

“The big promise of these neural scene representations, at the end of the day, is to use them in vision tasks. I give you an image and from that image you create a representation of the scene, and then everything you want to reason about you do in the space of that 3D scene,” says Vincent Sitzmann, a postdoc in the Computer Science and Artificial Intelligence Laboratory (CSAIL) and co-lead author of the paper.

Sitzmann wrote the paper with co-lead author Semon Rezchikov, a postdoc at Harvard University; William T. Freeman, the Thomas and Gerd Perkins Professor of Electrical Engineering and Computer Science and a member of CSAIL; Joshua B. Tenenbaum, a professor of computational cognitive science in the Department of Brain and Cognitive Sciences and a member of CSAIL; and senior author Frédo Durand, a professor of electrical engineering and computer science and a member of CSAIL. The research will be presented at the Conference on Neural Information Processing Systems this month.

Mapping rays

In computer vision and computer graphics, rendering a 3D scene from an image involves mapping thousands or possibly millions of camera rays. Think of camera rays like laser beams shooting out from a camera lens and striking each pixel in an image, one ray per pixel. These computer models must determine the color of the pixel struck by each camera ray.

Many current methods accomplish this by taking hundreds of samples along the length of each camera ray as it moves through space, which is a computationally expensive process that can lead to slow rendering.

Instead, an LFN learns to represent the light field of a 3D scene and then directly maps each camera ray in the light field to the color that is observed by that ray. An LFN leverages the unique properties of light fields, which enable the rendering of a ray after only a single evaluation, so the LFN doesn’t need to stop along the length of a ray to run calculations.

“With other methods, when you do this rendering, you have to follow the ray until you find the surface. You have to do thousands of samples, because that is what it means to find a surface. And you’re not even done yet because there may be complex things like transparency or reflections. With a light field, once you have reconstructed the light field, which is a complicated problem, rendering a single ray just takes a single sample of the representation, because the representation directly maps a ray to its color,” Sitzmann says.      

The LFN classifies each camera ray using its “Plücker coordinates,” which represent a line in 3D space based on its direction and how far it is from its point of origin. The system computes the Plücker coordinates of each camera ray at the point where it hits a pixel to render an image.

By mapping each ray using Plücker coordinates, the LFN is also able to compute the geometry of the scene due to the parallax effect. Parallax is the difference in apparent position of an object when viewed from two different lines of sight. For instance, if you move your head, objects that are farther away seem to move less than objects that are closer. The LFN can tell the depth of objects in a scene due to parallax, and uses this information to encode a scene’s geometry as well as its appearance.

But to reconstruct light fields, the neural network must first learn about the structures of light fields, so the researchers trained their model with many images of simple scenes of cars and chairs.

“There is an intrinsic geometry of light fields, which is what our model is trying to learn. You might worry that light fields of cars and chairs are so different that you can’t learn some commonality between them. But it turns out, if you add more kinds of objects, as long as there is some homogeneity, you get a better and better sense of how light fields of general objects look, so you can generalize about classes,” Rezchikov says.

Once the model learns the structure of a light field, it can render a 3D scene from only one image as an input.

Rapid rendering

The researchers tested their model by reconstructing 360-degree light fields of several simple scenes. They found that LFNs were able to render scenes at more than 500 frames per second, about three orders of magnitude faster than other methods. In addition, the 3D objects rendered by LFNs were often crisper than those generated by other models.

An LFN is also less memory-intensive, requiring only about 1.6 megabytes of storage, as opposed to 146 megabytes for a popular baseline method.

“Light fields were proposed before, but back then they were intractable. Now, with these techniques that we used in this paper, for the first time you can both represent these light fields and work with these light fields. It is an interesting convergence of the mathematical models and the neural network models that we have developed coming together in this application of representing scenes so machines can reason about them,” Sitzmann says.

In the future, the researchers would like to make their model more robust so it could be used effectively for complex, real-world scenes. One way to drive LFNs forward is to focus only on reconstructing certain patches of the light field, which could enable the model to run faster and perform better in real-world environments, Sitzmann says.

“Neural rendering has recently enabled photorealistic rendering and editing of images from only a sparse set of input views. Unfortunately, all existing techniques are computationally very expensive, preventing applications that require real-time processing, like video conferencing. This project takes a big step toward a new generation of computationally efficient and mathematically elegant neural rendering algorithms,” says Gordon Wetzstein, an associate professor of electrical engineering at Stanford University, who was not involved in this research. “I anticipate that it will have widespread applications, in computer graphics, computer vision, and beyond.”

This work is supported by the National Science Foundation, the Office of Naval Research, Mitsubishi, the Defense Advanced Research Projects Agency, and the Singapore Defense Science and Technology Agency.



de MIT News https://ift.tt/3GkozjJ

Making her way through MIT

Lucy Du, a doctoral student in the MIT Media Lab, has a remarkable passion for making. She spends her work day in lab designing and fabricating prosthetics, and devotes her free time to personal projects in the MIT MakerWorkshop or inspiring other students to try their hands at engineering. “The best feeling is when I get to go into a shop and make some parts, or order some parts — and the day they come in is like Christmas,” she says.

Her affinity for making started at a young age. “I loved building things and having tangible hardware to work on. Sitting there and coding or doing math all day was never what I wanted,” Du says. She participated in a robotics team in high school and has drawn inspiration from Disney movies and entertainment for many years, especially as their humanoid animatronic technology has grown. (Disney’s “eerily organic-looking” Spiderman stunt robot is a particular favorite of hers.)

Now, as a fourth-year PhD student, Du is channeling her passion for building things into designing a prosthetic ankle that is readily accessible to people of all sizes, since current commercial designs are only suited for tall individuals.

She also shares her love of engineering in her pursuits outside of the lab. Throughout her time here (Du also earned her undergraduate and master’s degree at MIT), she has found ways to make building things and engineering more accessible to others — from creating a student makerspace to teaching high school girls. She even turned a role on Discovery’s reality TV show “BattleBots” into an opportunity to inspire kids about engineering.

Making for prosthetics

When Du finished her master’s in mechanical engineering in 2016, she was ready to experience something outside of academia. She went on to work at NASA’s Jet Propulsion Laboratory, but after two years, she started to feel an itch to get her PhD. Even though she considered other schools, MIT stood out as an obvious choice. “MIT has so many resources and so many opportunities that you can be here for years and years and not even scratch the surface,” she says. “I think it was where I was meant to be.”

Du knew she wanted to work on animatronic robotics. Finding the right lab was not without bumps, but she ultimately ended up in the Biomechatronics Group under professor of media arts and sciences Hugh Herr. Housed within the Media Lab, it is an interdisciplinary lab broadly focused on prosthetics, exoskeletons, and the human-robot interface. She knew Herr’s lab was the right fit for her because of its emphasis on hardware design. “It is actually really hard to find robotics labs that focus on hardware building,” she notes. “A lot of robotics labs will buy the hardware for a project and focus solely on the software. I believe in designing your hardware with the end application in mind, as this can make the whole process better.”

To that end, Du’s research project focuses on designing the hardware for a robotic ankle that functions better than what is available now. “I hope the design will be able to fully mimic biological movements, including fast walking, walking up and down stairs and ramps, and some other common motions that you would make throughout a day,” she explains.

Currently, there is only one commercial powered prosthetic ankle that provides enough force for walking. But it has notable limitations, Du says. Because the design is large and bulky, “you either need to be a person of tall stature or you have to have a very short residual limb after amputation in order to wear the product.”

In contrast, her prosthetic ankle has a smaller profile that would enable more people to use it. She notes, “The design itself is meant to be scaled, so you can have the same design and scale it down for children or other people who don’t need as much power, and then scale up to larger adults.”

Inspiring other makers

Du has devoted considerable time and energy over her years at MIT to helping others explore making and engineering. As a master’s student, she served as an instructor for the Women’s Technology Program in Mechanical Engineering, a summer program for high school girls that aims to inspire them to pursue engineering. During a month-long crash course, she served as one of three graduate instructors for the mechanical engineering curriculum, teaching 20 high school girls the basics of kinematics and dynamics, and working with them on cool, hands-on experiments.

Mentoring, she says, is “probably the most rewarding thing that I’ve done” — especially watching some of those students attend MIT and excel. She has continued to cultivate her love of teaching and mentoring as a teaching assistant during her PhD program.

Du is also a founding member and leader of the MIT MakerWorkshop, the only completely student-run makerspace at MIT. As a master’s student, she noticed an unmet need for a makerspace where students could work on their personal projects, at hours that were convenient to them and did not conflict with class time. Even though there were already a lot of shops on campus, she says, “It was “pretty difficult to get access to a machine shop and for most of them, you were only supposed to work on class projects or research projects.”

The MakerWorkshop has served as much more than a workplace for Du; it has also been a hub for connections and inspiration throughout her graduate career. “A lot of times, I have an engineering problem or a life problem and I just want to talk to somebody. I could roam around the space, somebody would be there, and you could just talk to them about their experiences or have impromptu design reviews on the board.”

The connections that Du formed through MakerWorkshop led her in an unexpected direction: reality TV. She was one of a group of MakerWorkshop members who formed a team named SawBlaze for the Discovery show “BattleBots,” a revival of an old Comedy Central show from the early 2000s. In the show, teams build 250-pound robots to “fight to the death” in an arena. The SawBlaze team has competed in four seasons to date, starting in 2016 with season 2. “It is really different from the stuff we usually design and build [for research or class], because you are designing to material failure,” Du says. So far, the experience has taught her how to plan and design for the most extreme cases of impact, often relying on intuition, experience, and empirical testing, because these scenarios are usually beyond the limits of modeling. 

However, Du is less than interested in the screen time that the “BattleBots” role has brought her. Even though she is excited about the show’s upcoming season 6, she favors the outreach events, termed maker faires, where awe-struck kids excitedly point out their favorite robots and she has the opportunity to share how she got started in engineering.

This fall marks the 10th year of Du’s MIT career. As she begins to contemplate what she’ll do after she graduates, she’s keeping an open mind. She knows she wants to work on new technology development, whether that leads her to industry or academia. And she knows that teaching and mentoring will play an important role in her future. “The more time you put into teaching, the more rewarding it is,” she says. “To see your students get it and improve, that just means the world to me.”



de MIT News https://ift.tt/3ovegTP

Anantha Chandrakasan awarded 2022 IEEE Mildred Dresselhaus Medal

Anantha Chandrakasan, dean of the MIT School of Engineering and Vannevar Bush Professor of EECS, has been named the recipient of the 2022 IEEE Mildred Dresselhaus Medal. In the award citation, the IEEE noted Chandrakasan’s “contributions to ultralow-power circuits and systems, and leadership in academia and advancing diversity in the profession.”

Anantha Chandrakasan received BS, MS, and PhD degrees in electrical engineering and computer sciences from the University of California at Berkeley, in 1989, 1990, and 1994, respectively. He joined the MIT faculty in 1994. Additionally, Chandrakasan serves as co-chair of the MIT–IBM Watson AI Lab, the MIT-Takeda Program, and the MIT and Accenture Convergence Initiative for Industry and Technology, and chairs the MIT Climate and Sustainability Consortium.

He was the director of the MIT Microsystems Technology Laboratories from 2006 to 2011. From July 2011 through June 2017, he served as head of the Department of Electrical Engineering and Computer Science (EECS), during which time he spearheaded a number of initiatives that opened opportunities for students, postdocs, and faculty to conduct research, explore entrepreneurial projects, and engage with EECS. These programs include “SuperUROP,” a year-long independent research program that provides tools for students to do publication-quality research; the Rising Stars program, an annual event that convenes graduate and postdoc women for the purpose of sharing advice about the early stages of an academic career; and StartMIT, an independent activities period class that provides students and postdocs the opportunity to learn from and interact with industrial innovation leaders.

Chandrakasan has received awards including the 2009 Semiconductor Industry Association University Researcher Award, the 2013 IEEE Donald O. Pederson Award in Solid-State Circuits, an honorary doctorate from KU Leuven in 2016, the University of California at Berkeley EE Distinguished Alumni Award in 2017, and the 2019 IEEE Solid-State Circuits Society Distinguished Service Award. He was also recognized as the author with the highest number of publications in the 60-year history of the IEEE International Solid-State Circuits Conference (ISSCC), the foremost global forum for presentation of advances in solid-state circuits and systems-on-a-chip. In 2015, he was elected to the National Academy of Engineering, and in 2019 he was elected to the American Academy of Arts and Sciences. Chandrakasan is an ACM Fellow, as well as an IEEE Fellow who has served in various roles for the IEEE ISSCC, including program chair, Signal Processing Subcommittee chair, and Technology Directions Sub-committee chair; and he was the conference chair of ISSCC from 2010 to 2018. He serves as the senior technical advisor to the conference starting with ISSCC 2019.

Chandrakasan leads the MIT Energy-Efficient Circuits and Systems Group, whose research projects have addressed security hardware, energy harvesting, and wireless charging for the internet of things; energy-efficient circuits and systems for multimedia processing; and platforms for ultra-low-power biomedical electronics. He is a co-author of "Low Power Digital CMOS Design" (Kluwer Academic Publishers, 1995), "Digital Integrated Circuits" (Pearson Prentice-Hall, 2003, 2nd edition), and "Sub-threshold Design for Ultra-Low Power Systems" (Springer 2006). From 2016 to 2021, he served on the board of The Engine, an accelerator launched by MIT to support startup companies working on scientific and technological innovation with the potential for transformative societal impact. He currently serves on the board of Analog Devices, the SMART Governing Board, and the Board of Trustees of the Perkins School for the Blind.

Named in honor of the late MIT Institute professor and professor emerita of physics and electrical engineering Mildred Dresselhaus, whose innovations helped mold the history of advancements in science, technology, and education around the world, the IEEE Mildred Dresselhaus Medal is given annually in recognition of outstanding technical contributions in science and engineering that are of great impact to IEEE’s fields of interest.



de MIT News https://ift.tt/3xYlBi5

Feast or forage? Study finds circuit that helps a brain decide

MIT neuroscientists have discovered the elegant architecture of a fundamental decision-making brain circuit that allows a C. elegans worm to either forage for food or stop to feast when it finds a source. Capable of integrating multiple streams of sensory information, the circuit employs just a few key neurons to sustain long-lasting behaviors and yet flexibly switch between them as environmental conditions warrant.

“For a foraging worm, the decision to roam or to dwell is one that will strongly impact its survival.” says study senior author Steven Flavell, the Lister Brothers Career Development Associate Professor in the Picower Institute for Learning and Memory and the Department of Brain and Cognitive Sciences at MIT. “We thought that studying how the brain controls this crucial decision-making process could uncover fundamental circuit elements that may be deployed in many animals’ brains.”

This approach of studying simple invertebrates to gain basic insights into how the brain functions has a long tradition in neuroscience, Flavell says. For example, studies of how a squid nerve propagates electrical impulses led to the key insight explaining how brain cells fire in virtually all animals.

Though the critical component of brain circuitry identified by Flavell and colleagues may seem simple now that it has been revealed, finding it was anything but easy. Lead author Ni Ji, a postdoc in Flavell’s lab, used several advanced technologies, including one of the lab’s own inventions, to figure it out. The results of her and her co-authors’ work appear in the journal eLife.

Tracking thinking

C. elegans is a popular model in neuroscience because it only has 302 neurons and the “wiring diagram,” or connectome, has been fully mapped. But even so, the very dense and overlapping interconnectedness among those neurons, plus their ability to signal each other via chemicals called neuromodulators, means that one can hardly just look at the connectome and discern how it switches between different states of behavior.

To identify functional circuitry amid this web of connections, Flavell’s lab developed a new microscope capable of tracking the worms as they move around, thereby constantly imaging the activity of neurons across the worm’s brain, as indicated by calcium-triggered flashes of light. Ji used the scope to focus on 10 interconnected neurons involved in foraging, tracking their patterns of neural activity associated with roaming or dwelling behaviors.

Ji and co-authors trained software that learned the patterns so well that just based on neural activity, it could predict the worm’s behavior with 95 percent accuracy. The analysis revealed a quartet of neurons whose activity was specifically associated with roaming. Another key pattern was that the transition from roaming around to stopping to dwell always followed activation of a neuron called NSM. Flavell’s lab previously showed that NSM can sense the presence of newly ingested food and emit a neuromodulator called serotonin to signal other neurons to slow the worm down to dwell in a nutritive area.

Mutual antagonism

Having identified the activity patterns that changed as the worm switched states, Ji began manipulating neurons in the circuit to understand how they interact. To confirm NSM’s role as the trigger of the dwelling state, Ji engineered it to be artificially activated with a flash of light (a technique called optogenetics). When she flashed the light, it caused the worm to dwell by inhibiting the activity of the roaming-associated neurons. Further experiments showed that this inhibitory power depended on the roaming neurons having an inhibitory serotonin receptor, called MOD-1. If Ji genetically knocked out the MOD-1 receptor, NSM couldn’t inhibit the roaming behavior and quickly stopped trying for lack of feedback.

Similarly, Ji showed that when the worm was roaming, it was because the roaming quartet was using the neuromodulator PDF to inhibit the activity of NSM. Optogenetic activation of PDF-expressing neurons tamped down NSM activity, for instance.

In a normal worm, if the roaming quartet was active NSM was not, and vice versa. But when Ji genetically knocked out the circuit elements that underlie this mutual inhibition both the roaming quartet and NSM could be active at the same time, leaving the worm in a weird state of meandering around at about half of roaming speed.

Sensory inputs

Thus, via an ongoing battle of mutual inhibition, roaming is sustained by the quartet and dwelling is sustained by NSM, but that still begged the question: How does the worm decide to flip the switch? To find out, Ji and colleagues programmed a machine learning algorithm to discern which neurons might work upstream in the broader circuit to influence the serotonin and PDF tug of war. This approach identified a neuron called AIA, which is known for integrating sensory information about food odors. AIA’s activity co-varied with a couple of the roaming neurons during roaming, and with NSM when dwelling began.

In other words, upon becoming activated by the smell of food, AIA could use its input to drive either side of the mutual inhibitory circuit to switch behavior. Remembering that NSM can sense when the worm is actually eating, Ji and Flavell could deduce what AIA and NSM must be doing. If the worm smells food but is not eating, it needs to roam further to that food smell until it is. If the worm smells food and at the same time it begins eating, then it should continue to dwell there.

“To a foraging worm, food odors are an important, but ambiguous, sensory cue. AIA’s ability to detect food odors and to transmit that information to these different downstream circuits, dependent on other incoming cues, allows animals to contextualize the smell and make adaptive foraging decisions,” Flavell said. “If you are looking for circuit elements that could also be operating in bigger brains, this one stands out as a basic motif that might allow for context-dependent behaviors.”

In addition to Ji and Flavell, the paper’s other authors are Gurrein Madan, Guadalupe Fabre, Alyssa Dayan, Casey Baker, Talya Kramer, and Ijeoma Nwabudike.

Funding for the research came from the National Institutes of Health, the National Science Foundation, the JPB Foundation, the Brain and Behavior Research Foundation, NARSAD, the McKnight Foundation, and the Alfred P. Sloan Foundation.



de MIT News https://ift.tt/31y1T0j

Generating a realistic 3D world

While standing in a kitchen, you push some metal bowls across the counter into the sink with a clang, and drape a towel over the back of a chair. In another room, it sounds like some precariously stacked wooden blocks fell over, and there’s an epic toy car crash. These interactions with our environment are just some of what humans experience on a daily basis at home, but while this world may seem real, it isn’t.

A new study from researchers at MIT, the MIT-IBM Watson AI Lab, Harvard University, and Stanford University is enabling a rich virtual world, very much like stepping into “The Matrix.” Their platform, called ThreeDWorld (TDW), simulates high-fidelity audio and visual environments, both indoor and outdoor, and allows users, objects, and mobile agents to interact like they would in real life and according to the laws of physics. Object orientations, physical characteristics, and velocities are calculated and executed for fluids, soft bodies, and rigid objects as interactions occur, producing accurate collisions and impact sounds.

TDW is unique in that it is designed to be flexible and generalizable, generating synthetic photo-realistic scenes and audio rendering in real time, which can be compiled into audio-visual datasets, modified through interactions within the scene, and adapted for human and neural network learning and prediction tests. Different types of robotic agents and avatars can also be spawned within the controlled simulation to perform, say, task planning and execution. And using virtual reality (VR), human attention and play behavior within the space can provide real-world data, for example.

“We are trying to build a general-purpose simulation platform that mimics the interactive richness of the real world for a variety of AI applications,” says study lead author Chuang Gan, MIT-IBM Watson AI Lab research scientist.

Creating realistic virtual worlds with which to investigate human behaviors and train robots has been a dream of AI and cognitive science researchers. “Most of AI right now is based on supervised learning, which relies on huge datasets of human-annotated images or sounds,” says Josh McDermott, associate professor in the Department of Brain and Cognitive Sciences (BCS) and an MIT-IBM Watson AI Lab project lead. These descriptions are expensive to compile, creating a bottleneck for research. And for physical properties of objects, like mass, which isn’t always readily apparent to human observers, labels may not be available at all. A simulator like TDW skirts this problem by generating scenes where all the parameters and annotations are known. Many competing simulations were motivated by this concern but were designed for specific applications; through its flexibility, TDW is intended to enable many applications that are poorly suited to other platforms.

Another advantage of TDW, McDermott notes, is that it provides a controlled setting for understanding the learning process and facilitating the improvement of AI robots. Robotic systems, which rely on trial and error, can be taught in an environment where they cannot cause physical harm. In addition, “many of us are excited about the doors that these sorts of virtual worlds open for doing experiments on humans to understand human perception and cognition. There's the possibility of creating these very rich sensory scenarios, where you still have total control and complete knowledge of what is happening in the environment.”

McDermott, Gan, and their colleagues are presenting this research at the conference on Neural Information Processing Systems (NeurIPS) in December.

Behind the framework

The work began as a collaboration between a group of MIT professors along with Stanford and IBM researchers, tethered by individual research interests into hearing, vision, cognition, and perceptual intelligence. TDW brought these together in one platform. “We were all interested in the idea of building a virtual world for the purpose of training AI systems that we could actually use as models of the brain,” says McDermott, who studies human and machine hearing. “So, we thought that this sort of environment, where you can have objects that will interact with each other and then render realistic sensory data from them, would be a valuable way to start to study that.”

To achieve this, the researchers built TDW on a video game platform called Unity3D Engine and committed to incorporating both visual and auditory data rendering without any animation. The simulation consists of two components: the build, which renders images, synthesizes audio, and runs physics simulations; and the controller, which is a Python-based interface where the user sends commands to the build. Researchers construct and populate a scene by pulling from an extensive 3D model library of objects, like furniture pieces, animals, and vehicles. These models respond accurately to lighting changes, and their material composition and orientation in the scene dictate their physical behaviors in the space. Dynamic lighting models accurately simulate scene illumination, causing shadows and dimming that correspond to the appropriate time of day and sun angle. The team has also created furnished virtual floor plans that researchers can fill with agents and avatars. To synthesize true-to-life audio, TDW uses generative models of impact sounds that are triggered by collisions or other object interactions within the simulation. TDW also simulates noise attenuation and reverberation in accordance with the geometry of the space and the objects in it.

Two physics engines in TDW power deformations and reactions between interacting objects — one for rigid bodies, and another for soft objects and fluids. TDW performs instantaneous calculations regarding mass, volume, and density, as well as any friction or other forces acting upon the materials. This allows machine learning models to learn about how objects with different physical properties would behave together.

Users, agents, and avatars can bring the scenes to life in several ways. A researcher could directly apply a force to an object through controller commands, which could literally set a virtual ball in motion. Avatars can be empowered to act or behave in a certain way within the space — e.g., with articulated limbs capable of performing task experiments. Lastly, VR head and handsets can allow users to interact with the virtual environment, potentially to generate human behavioral data that machine learning models could learn from.

Richer AI experiences

To trial and demonstrate TDW’s unique features, capabilities, and applications, the team ran a battery of tests comparing datasets generated by TDW and other virtual simulations. The team found that neural networks trained on scene image snapshots with randomly placed camera angles from TDW outperformed other simulations’ snapshots in image classification tests and neared that of systems trained on real-world images. The researchers also generated and trained a material classification model on audio clips of small objects dropping onto surfaces in TDW and asked it to identify the types of materials that were interacting. They found that TDW produced significant gains over its competitor. Additional object-drop testing with neural networks trained on TDW revealed that the combination of audio and vision together is the best way to identify the physical properties of objects, motivating further study of audio-visual integration.

TDW is proving particularly useful for designing and testing systems that understand how the physical events in a scene will evolve over time. This includes facilitating benchmarks of how well a model or algorithm makes physical predictions of, for instance, the stability of stacks of objects, or the motion of objects following a collision — humans learn many of these concepts as children, but many machines need to demonstrate this capacity to be useful in the real world. TDW has also enabled comparisons of human curiosity and prediction against those of machine agents designed to evaluate social interactions within different scenarios.

Gan points out that these applications are only the tip of the iceberg. By expanding the physical simulation capabilities of TDW to depict the real world more accurately, “we are trying to create new benchmarks to advance AI technologies, and to use these benchmarks to open up many new problems that until now have been difficult to study.”

The research team on the paper also includes MIT engineers Jeremy Schwartz and Seth Alter, who are instrumental to the operation of TDW; BCS professors James DiCarlo and Joshua Tenenbaum; graduate students Aidan Curtis and Martin Schrimpf; and former postdocs James Traer (now an assistant professor at the University of Iowa) and Jonas Kubilius PhD ‘08. Their colleagues are IBM director of the MIT-IBM Watson AI Lab David Cox; research software engineer Abhishek Bhandwalder; and research staff member Dan Gutfreund of IBM. Additional researchers co-authoring are Harvard University assistant professor Julian De Freitas; and from Stanford University, assistant professors Daniel L.K. Yamins (a TDW founder) and Nick Haber, postdoc Daniel M. Bear, and graduate students Megumi Sano, Kuno Kim, Elias Wang, Damian Mrowca, Kevin Feigelis, and Michael Lingelbach.

This research was supported by the MIT-IBM Watson AI Lab.



de MIT News https://ift.tt/3ppl1pC

Jelena Vučković delivers 2021 Dresselhaus Lecture on inverse-designed photonics

As her topic for the 2021 Mildred S. Dresselhaus Lecture, Stanford University professor Jelena Vučković posed a question: Are computers better than humans in designing photonics?

Throughout her talk, presented on Nov. 15 in a hybrid format to more than 500 attendees, the Jensen Huang Professor in Global Leadership at Stanford’s School of Engineering offered multiple examples arguing that, yes, computer software can help identify better solutions than traditional methods, leading to smaller, more efficient devices, as well as entirely new functionalities.

Photonics, the science of guiding and manipulating light, is used in many applications such as optical interconnects, optical computing platforms for AI or quantum computing, augmented reality glasses, biosensors, medical imaging systems, and sensors in autonomous vehicles.

For all these applications, Vučković said, many optical components must be integrated on a chip that can fit into the footprint of your glasses or mobile device. Unfortunately, the problems with high-density photonic integration are several. Traditional photonic components are large, sensitive to fabrication errors and environmental factors such as variations in temperature, and are designed by manual tuning with few parameters. So, Vučković and her team asked, “How can we design better photonics?”

Her answer: photonics inverse design. In this process, scientists rely on sophisticated computational tools and modern computing platforms to discover optimal photonic solutions or device designs for a particular function. In this inverse process, the researcher first considers how he or she would like the photonic block to operate, then uses computer software to search the whole parameter space of possible solutions for the one that is optimal, within fabrication restraints.

From guiding light around corners to splitting colors of light in a compact footprint, Vučković presented several examples to prove this process works — using computer software to conduct physics-guided searches of numerous possibilities produces non-traditional solutions that increase efficiency and/or decrease the footprint of photonic devices.

Enabling new functionalities – high-energy physics

State-of-the-art particle accelerators, which use microwaves or radio frequency waves to propel charged particles, can be the size of a full city block; Stanford’s SLAC National Accelerator Lab, for example, is two miles long. Lower-energy accelerators, such as those used in medical radiation facilities, are not as large, but still take up an entire room, are expensive, and not very accessible. “If we could use a different spectrum of electromagnetic waves with shorter wavelengths to do the same function of accelerating particles,” said Vučković, “we should, in principle, be able to shrink the size of an accelerator.” The solution is not as simple as reducing the size of all the parts, as the electromagnetic building blocks won’t work for optical waves. Instead, Vučković and her team used the inverse design process to create new building blocks, and built a single-stage on-chip integrated laser-driver particle accelerator that is only 30 micrometers in length.

Applying inverse designed photonics to practical environments

Autonomous vehicles have a large lidar system on the roof housing mechanics that enable rotation of a beam to scan the environment. Vučković considers how this could be improved. “Can you make this system inside the footprint of a single chip, which would be just like another sensor in your car, and can it be inexpensive?” Through inverse design, her research group found optimal photonic structures to enable steering the beam with inexpensive lasers that are cheaper, and achieve 5 degrees of additional beam steering, than state-of-the-art systems.

Next up: scaling superconducting quantum processors onto a single diamond or silicon carbide chip. In this example, Vučković harkened back to the 2020 Dresselhaus Lecture delivered by Harvard Professor Evelyn Hu on leveraging defects at the nanoscale. By relying on impurities in these materials at low concentrations, naturally trapped atoms can be very useful for quantum applications. Vučković’s group is working on material developments and fabrication techniques that allow them to put these trapped atoms in desired positions with minimal defects.

“For many applications, letting computer software search for an optimal solution leads to better solutions than what you would design, or guess, based on your intuitions. And this process is material-agnostic, fully compatible with commercial foundries, and enables new functionalities,” said Vučković. “Even if you try to make something a little bit better than traditional solutions — smaller in a footprint or higher in efficiency — we can come up with multiple solutions that are equally good or better than what we knew before. We are relearning photonics and electromagnetics in this process.”

Honoring Mildred S. Dresselhaus and Gene Dresselhaus

Vučković was the third speaker to deliver the Dresselhaus Lecture, established in 2019 to honor the late MIT physics and electrical engineering professor Mildred Dresselhaus. This year, the lecture was also dedicated to Gene Dresselhaus, renowned physicist and Millie’s husband, who passed away in late September 2021.

Jing Kong, professor of electrical engineering and computer science at MIT, opened the lecture by reflecting on the Dresselhaus’s scientific achievements. Kong highlighted the American Physical Society Oliver E Buckley Condensed Matter Physics Prize — considered the most prestigious award granted within the field of condensed-matter physics — that was awarded to both Millie (2008) and Gene (2022). “Although they worked together on many important topics,” said Kong, “It’s remarkable that they received this award for separate research works. It is humbling for us to walk in their footsteps.”



de MIT News https://ift.tt/3ImZPcp