Just in:
Bitcoin steadies as gold extends powerful rally // MoreTickets Reveals Hong Kong’s Top-Searched Summer Events and Evolving Ticket-Buying Behaviours // Wuxi Symphony Orchestra Debuts at Ljubljana Festival: Sounds of the East Illuminate the Historic Central European City // Grok glitch sends users streams of gibberish // Jharkhand may witness a more complicated students movement // Standard Chartered takes HKDAP into banking mainstream // Iran rial sinks beyond two million per dollar // Silence Wang Wax Figure Arrives at Madame Tussauds Hong Kong // Iran braces for sweeping US economic offensive // North Korea-linked hackers target Rust software supply chain // From Vietnam to the U.S: East West Barbershop takes on the world’s most competitive market // ADNOC Distribution brings Reatile into Shell deal // Six Leading Enterprises Jointly Awarded Tender for Hung Shui Kiu/Ha Tsuen Pilot Development Area // TDCX Opens Second Hyderabad Campus, Reinforcing India as Key Global Delivery Hub // Treasury weighs cash account for expanded bond buybacks // Anonymous Ox Alpha raises questions over prompt retention // 40 Teams Gather in Hong Kong to Compete in the “AI x Cybersecurity Challenge” // MyRepublic expands GAMER lineup with Dreamcore x MyRepublic RTX 5060 Ti Gaming PC and Limited Edition ASUS T1 Graphics Card Broadband Bundle // TDCX expands Hyderabad footprint with second campus // 5G Capital Sets a New Benchmark:China Unicom Beijing and Huawei Power the 2nd World Humanoid Robot Games with 5G-A GigaUplink //

MIT's new AI: So smart it can predict what happens next from a still image

1480430109 mitvideoainov16

mitvideoainov16.jpg

The AI can add animations to static images, although the researchers acknowledge that it rarely generates the “correct” future.


Image: MIT

Researchers have developed a deep-learning system that can do the very human task of interpreting what’s happening in a photo and guessing what’s likely to happen next.

Better yet, the system, developed by machine-learning researchers at MIT, can express its idea of a plausible future by adding animations to still images, such as waves that would ultimately crash, people who might move in a field, or a train that might roll forward on its tracks.

The work could provide a new direction for exploring computer vision by giving machines the ability to understand how objects move in the real world.

The researchers achieved their objective by training two deep networks on thousands of hours of unlabelled video. As the researchers note, annotating video is expensive, but there’s no shortage of unlabeled video to begin training machines to read other signals about the world.

Carl Vondrick, a PhD student at MIT, who specializes in machine learning and computer vision, told New Scientist that the ability to predict movements in a scene could ensure tomorrow’s domestic helper robots don’t become a hindrance. You wouldn’t, for example, appreciate a robot pulling a chair out from under you as you’re about to sit down, he said.

The model was trained on two million videos from Flickr amounting to 5,000 hours of content covering four main scene types, including golf courses, beaches, train stations, and hospital rooms, consisting mostly of images of babies.

As New Scientist notes, the videos the model produces are grainy and short, lasting about one second, but they do capture the right general movement of a given scene, such as a train moving forward or a baby scrunching up its face.

However, the model still has much to learn about how the world works, such as that a train doesn’t infinitely depart from a scene. Still, it did show that machines can be taught to dream up brief plausible futures.

The model is also able to “hallucinate” fictional but reasonable motions for each of the scene categories they explored.

The model was based on a machine-learning technique called adversarial learning, where two deep networks compete against each other. One network generates synthetic video while the other tries to discriminate between generated video and real videos.

Vondrick has previously also trained deep-learning models on hundreds of hours of unlabeled YouTube videos and TV programs such as ‘The Office’ to predict human interactions and gestures, such as a handshake, hug or kiss.

Read more research from MIT

(via PCMag)



Notice an issue?

Arabian Post strives to deliver the most accurate and reliable information to its readers. If you believe you have identified an error or inconsistency in this article, please don't hesitate to contact our editorial team at editor[at]thearabianpost[dot]com. We are committed to promptly addressing any concerns and ensuring the highest level of journalistic integrity.


Loading next story…