Rarely do I miss an opportunity to apply techniques from work to things I enjoy doing in my free time. Often that thing is cycling, the greatest sport in the world. But when autumn comes around and the streets turn cold and dark, I prefer the iron of the weight room over the carbon of the road. A technical sport like Olympic weightlifting is of course best mastered with the guidance of an experienced coach. To the less fortunate or ambitious, video analysis may help hone skills and track progress. Various apps offer basic video tools, but I’ve yet to find one that provides the data I am truly interested in: ground reaction forces (GRFs), forces applied to the barbell, accurate kinematics, and mechanical power output. That is why I developed my own weightlifting video analyzer. The demo below is the result of several techniques from work combined: AI & computer vision (Raidyn), state-input estimation & multibody dynamics (KU Leuven Mecha(tro)nic System Dynamics (LMSD)), and biomechanical impact modelling (Classified Cycling). At the core is a flexible multibody model of the barbell that feeds into a combined state-input-parameter estimator to infer the forces applied to the bar. With accurate barbell forces, GRFs can be estimated with higher precision than methods based solely on accelerations derived from video. In addition, the model unlocks a virtually unlimited supply of synthetic training data for the AI model, ensuring robust segmentation of barbell motion and deformation. Initially I cast myself as the hero of the demo, but I soon realized that a video starring the Pogačar of weightlifting would make for a more spectacular analysis. Enter Bulgarian Karlos Nasar, 20 years of age at this year’s European Championships in Moldova, lifting a record-breaking 229 kg overhead. Running my analyzer on Nasar’s low-resolution YouTube video allowed me to test its performance on a subject it hadn’t seen before in training, with frame rate and resolution far below what I use in my own sessions – and yet the algorithms didn’t miss a beat. Sure, it helps that weightlifting consists of only two precisely defined movement patterns, but I was still pleasantly surprised. The demo below is just the start. Belgium’s cold season still has a way to go and I’ve got plenty of ideas to improve and expand the algorithms. Next up: joint and muscle force estimation.
Advanced Computer Vision Techniques
Explore top LinkedIn content from expert professionals.
-
-
Presenting FEELTHEFORCE (FTF): a robot learning system that models human tactile behavior to learn force-sensitive manipulation. Using a tactile glove to measure contact forces and a vision-based model to estimate hand pose, they train a closed-loop policy that continuously predicts the forces needed for manipulation. This policy is re-targeted to a Franka Panda robot with tactile gripper sensors using shared visual and action representa- tions. At execution, a PD controller modulates gripper closure to track predicted forces -enabling precise, force-aware control. This approach grounds robust low- level force control in scalable human supervision, achieving a 77% success rate across 5 force-sensitive manipulation tasks. #research: https://lnkd.in/dXxX7Enw #github: https://lnkd.in/dQVuYTDJ #authors: Ademi Adeniji, Zhuoran (Jolia) Chen, Vincent Liu, Venkatesh Pattabiraman, Raunaq Bhirangi, Pieter Abbeel, Lerrel Pinto, Siddhant Haldar New York University, University of California, Berkeley, NYU Shanghai Controlling fine-grained forces during manipulation remains a core challenge in robotics. While robot policies learned from robot-collected data or simulation show promise, they struggle to generalize across the diverse range of real-world interactions. Learning directly from humans offers a scalable solution, enabling demonstrators to perform skills in their natural embodiment and in everyday environments. However, visual demonstrations alone lack the information needed to infer precise contact forces.
-
Robots on the pitch....You better believe it. Will you be able to play with this one? No more standing cones or passive drills. Athletes today are dodging dynamic robots—machines that track, move, and react in real time. These aren’t gimmicks; they’re next-gen training partners. ⚽ In football, systems like SKILLSLAB, Rezzil, and Trailblazer Training Bots are already used by top clubs to simulate high-pressure situations, improve decision-making, and measure milliseconds of reaction time. 🏀 In basketball, robotic arms help perfect shooting arcs, while AI vision tools break down footwork frame by frame. 🎾 In tennis, smart ball machines adjust spin, speed, and placement in unpredictable sequences—training the brain as much as the body. Why it matters: + Athletes improve reaction speed by up to 20% using adaptive robotic drills. + Training bots allow 3x more touches per minute compared to traditional drills. + Machine-learning platforms track thousands of data points per session—customizing feedback instantly. This isn’t just tech—it’s transformation. Robots are helping players train faster, smarter, and with a grin on their face. #Innovation #Tech #Robots
-
Streaming 3D reconstruction is fundamentally a memory problem. How do you map a massive, multi-room environment without blowing up your compute budget as the sequence gets longer? Lingbo-Map just introduced a highly elegant architectural solution to this exact bottleneck: Geometric Context Attention (GCA). Instead of brute-forcing the entire scene history into memory, GCA splits the streaming state into three lightweight buckets: an anchor for global coordinate grounding, a local reference window for dense geometry, and a compressed trajectory memory. By squashing the full sequence history into compact per-frame tokens, the memory and compute requirements remain nearly constant. Running through a DINO backbone, the pipeline actively predicts camera poses and depth maps at ~20 FPS—even on continuous 10,000+ frame sequences. This is how you scale real-time spatial computing and large-scale digital twins without needing infinite VRAM. Models: https://lnkd.in/dxY7D4Ar Project page: https://lnkd.in/dKRUEQaq Code: https://lnkd.in/dXQSJB7u Paper: https://lnkd.in/diPQk3Ki #SpatialComputing #3DReconstruction #ComputerVision #MachineLearning #SLAM #DevRel
-
Introducing SAM 3D: Powerful 3D Reconstruction for Physical World Images and it’s not your typical 3D reconstruction tool. It does what previous models couldn’t: - Reconstruct real-world objects and scenes from a single image. - Handle occlusion, indirect views, and cluttered backgrounds. - Estimate human pose and shape with surprising accuracy. Why does this matter? Because for the first time, we’re seeing 3D perception at the scale, quality, and accessibility of today’s 2D models. The architecture behind SAM 3D borrows from LLMs: pre-training on synthetic data, followed by post-training on real-world images using a human-in-the-loop ranking engine. The result is a feedback loop that continuously improves both the data and the model. The implications stretch far beyond creative media. Robotics. E-commerce. Sports medicine. Interactive avatars. You name it. And the best part? It's fast. Real-time fast. Meta’s already using SAM 3D to power a “View in Room” feature in Facebook Marketplace; turning static listings into immersive experiences. The gap between virtual and physical is shrinking. SAM 3D is a serious leap forward. Learn more: - 𝗥𝗲𝗮𝗱 𝘁𝗵𝗲 𝗮𝗻𝗻𝗼𝘂𝗻𝗰𝗲𝗺𝗲𝗻𝘁 𝗯𝗹𝗼𝗴: https://lnkd.in/g8w6dvAB - 𝗥𝗲𝗮𝗱 𝘁𝗵𝗲 𝗦𝗔𝗠 𝟯𝗗 𝗢𝗯𝗷𝗲𝗰𝘁𝘀 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵 𝗣𝗮𝗽𝗲𝗿: https://lnkd.in/gd_HQE9c - 𝗥𝗲𝗮𝗱 𝘁𝗵𝗲 𝗦𝗔𝗠 𝟯𝗗 𝗕𝗼𝗱𝘆 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵 𝗣𝗮𝗽𝗲𝗿: https://lnkd.in/gp3p9jaf - 𝗗𝗼𝘄𝗻𝗹𝗼𝗮𝗱 𝗦𝗔𝗠 𝟯𝗗 𝗢𝗯𝗷𝗲𝗰𝘁𝘀: https://lnkd.in/gfwmcKGK - 𝗗𝗼𝘄𝗻𝗹𝗼𝗮𝗱 𝗦𝗔𝗠 𝟯𝗗 𝗕𝗼𝗱𝘆: https://lnkd.in/g9CPHEXJ - 𝗘𝘅𝗽𝗹𝗼𝗿𝗲 𝘁𝗵𝗲 𝗣𝗹𝗮𝘆𝗴𝗿𝗼𝘂𝗻𝗱: https://lnkd.in/guSDTnU3
-
Meta just introduced #SAM-3D, a model that can turn a single photo into a complete 3D scene - geometry, texture, pose, and even hidden structure. It works on real, cluttered images where previous models failed. Why it’s a breakthrough: SAM-3D doesn’t just fill in visible pixels. It reconstructs the full 3D shape and places objects correctly in the scene. This is the closest step yet toward a general 3D foundation model. How Meta achieved it: - A hybrid “human + model-in-the-loop” pipeline - Nearly 1M real images - 3.14M meshes - LLM-style pretrain → mid-train → post-train → DPO alignment Performance gains: - 5× human preference wins on object reconstructions - 6× win on full scenes - Best-in-class Chamfer distance (0.0400) - Geometry inference reduced from 25 steps to 4 Why it matters: - This raises the bar for robotics, AR/VR, gaming, advertising, and any workflow that needs fast, accurate 3D. With SAM-3D, Meta is positioning itself at the front of spatial AI. #AI #3DReconstruction #ComputerVision #SpatialAI #GenerativeAI #DeepLearning #AR #VR #Robotics #MetaAI
-
This week's defining shift for me is that creating 3D data is getting much simpler. New tools are turning everyday inputs like smartphone video, single photos, and text prompts into usable 3D environments and assets. This lowers the barrier to building the scenes, objects, and spaces that robotics, simulation, and immersive content rely on. It also shifts 3D creation from a specialized skill to something all teams can generate quickly and at the scale modern spatial systems require. This week’s news surfaced signals like these: 🤖 Parallax Worlds raised $4.9 million to turn standard video into digital twins for robotics testing. The platform turns basic walkthrough videos into interactive 3D spaces that teams can use to run their robot software and see how it performs before sending anything into the field. 🪑 Meta introduced SAM 3D to reconstruct objects and people from single images, producing full-textured meshes even when subjects are partly hidden or shot from difficult angles. The models were trained using real-world data and a staged process to improve accuracy. 🌏 Meta unveiled WorldGen, a research tool that generates full 3D worlds from text prompts. It produces complete, navigable spaces that can be used in Unity or Unreal and shows how AI can create environments without manual modeling. Why this matters: Faster 3D pipelines expand who can build, test, and refine spatial ideas. They turn 3D creation from a bottleneck into a regular part of development, which opens the door to more experimentation and better decisions earlier in the process. #robotics #digitaltwins #simulation #VR #AR #virtualreality #spatialcomputing #physicalAI #AI #3D
-
New YouTube tutorial: Football AI step-by-step guide! ⚽ ⚽ ⚽ I poured 200+ hours into this, and it's finally here! In just 1.5 hours, you'll learn how to: - detecting balls, players, goalkeepers, and referees, including fine-tuning on a custom dataset - tracking players in real-time - assigning teams using SigLIP and embedding analysis - identifying 32 key points on the football field with custom dataset fine-tuning - performing homography perspective transformation with OpenCV - implementing advanced analytics, such as radar view, Voronoi diagrams, and ball trajectory mapping 🔥 🔥 🔥 ⮑ 🔗 tutorial: https://lnkd.in/d4RY4-Zu Find links to related resources and cool visualizations in the comments below 👇🏻 #computervision #objectdetection #sportsanalysis #opensource
-
𝗧𝗵𝗲 𝗱𝗿𝗼𝗻𝗲’𝘀 𝘃𝗶𝗱𝗲𝗼 𝗳𝗲𝗲𝗱 𝗰𝗮𝗻 𝗻𝗼𝘄 𝗯𝗲𝘁𝗿𝗮𝘆 𝗶𝘁𝘀 𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻. Ukrainian-Estonian defence-tech company Farsight Vision has unveiled FSV Localizer, software designed to geolocate an enemy UAV from intercepted footage or a live onboard video stream. The operator selects an approximate search area. The software then analyses individual frames, matches visible terrain against geospatial data and plots the drone’s estimated positions and direction of travel on a map. The company says this can reveal where the drone came from, where it is heading and approximately where it is now—in seconds and without dedicated radio direction-finding equipment. That is potentially significant. An intercepted video link has traditionally provided insight into what the enemy operator sees. It may now also expose the aircraft’s route, target approach and, in favourable circumstances, clues to its launch area or control team. The previously intercepted footage from a Russian Molniya strike drone is exactly the type of material such a system could exploit: roads, buildings, fields, power infrastructure and other persistent terrain features pass beneath the camera throughout the flight. But this does not make RF direction finding obsolete. The stream must first be intercepted, the broad search area must be known, and the imagery must contain terrain distinctive enough to match reliably. Darkness, smoke, cloud, featureless ground and outdated reference data will all complicate the result. The real value is therefore fusion. Combine visual geolocation with RF detection, radar tracks and acoustic or optical sensors, and every intercepted frame becomes another measurement in the counter-UAS picture. 𝘛𝘩𝘦 𝘷𝘪𝘥𝘦𝘰 𝘧𝘦𝘦𝘥 𝘪𝘴 𝘯𝘰 𝘭𝘰𝘯𝘨𝘦𝘳 𝘫𝘶𝘴𝘵 𝘵𝘩𝘦 𝘥𝘳𝘰𝘯𝘦’𝘴 𝘦𝘺𝘦𝘴. 𝘐𝘵 𝘪𝘴 𝘢𝘭𝘴𝘰 𝘢 𝘵𝘳𝘢𝘪𝘭 𝘣𝘢𝘤𝘬 𝘵𝘰 𝘵𝘩𝘦 𝘥𝘳𝘰𝘯𝘦.
-
For robot dexterity, a missing piece is general, robust perception. Our new Science Robotics article combines multimodal sensing with neural representations to perceive novel objects in-hand. See it on the cover of the November issue! https://lnkd.in/ezZRs5dN We estimate pose and shape by learning neural field models online from a stream of vision, touch, and proprioception. The frontend achieves robust segmentation and depth prediction for vision and touch. The backend combines this information into a neural field, while also optimizing for pose. Vision-based touch (digit.ml/digit) perceives contact geometries as images, and we train an image-to-depth tactile transformer in simulation. For visual segmentation, we combine powerful foundation models (SAMv1) with robot kinematics. It doubles up as a multimodal pose tracker, when provided CAD models of the objects at runtime. For different levels of occlusion, we find that “touch, at the very least, refines and, at the very best, disambiguates visual estimates during in-hand manipulation." We release a large dataset of real-world and simulated visuo-tactile interactions and tactile transformer models on Hugging Face: bit.ly/hf-neuralfeels This has been in the pipeline for a while, thanks to my amazing collaborators from AI at Meta, Carnegie Mellon University, University of California, Berkeley, Technische Universität Dresden, and CeTI: Haozhi Qi, Tingfan Wu, Taosha Fan, Luis Pineda, Mike Maroje Lambeta, Jitendra MALIK, Mrinal Kalakrishnan, Roberto Calandra, Michael Kaess, Joseph Ortiz, and Mustafa Mukadam Paper: https://lnkd.in/ezZRs5dN Project page: https://lnkd.in/dCPCs4jQ #ScienceRoboticsResearch