Skip to content
View bowieshi's full-sized avatar

Highlights

  • Pro

Block or report bowieshi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bowieshi/README.md

Hi there 👋

""I am quite interested about R&D high-performance, high-impact and robust machine learning system which survives the real world.""

My name is Shi Boao. My English name is Bowie. I am currently a research assistant at Penn GRASP Lab. I am also a final year undergraduate majoring in Computer Science from The University of Hong Kong.

I am contributing to VLLM-OMNI now. I grow quite interested about building high-performance, production-ready machine learning systems.

My research interests lie in General 3D/4D, Physical intelligence and Robotics. I want to develop scalable perceptual and physical world modelling systems that can learn directly from large scale unconstrained video, arguably the most abundant and accessible form of real-world data. With this in hand, how can we enable robotics agent decision making training by simulation in scale.

Several directions/topics I am interested in:

  • Perception foundation model (Visual geometry model, Feed forward 3D/4D reconstruction)
  • 4D foundation model (Action conditioned world model)
  • Robotics learning

My current research goal is to answer the following questions:

  • How can we learn dynamics and interaction causality of the real world from large scale streaming monocular observation?
  • How can we reason about the underlying intrinsic or physical property of unstructured entity representation (pixels, point clouds, tracklets, etc.)?
  • How can we extract scalable and generalizable priors of world dynamics and interaction causality from human interaction and decision videos, and how to exploit it as a foundation for robotics tasks?

Pinned Loading

  1. vllm-project/vllm-omni vllm-project/vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 6k 1.4k

  2. UCB_CS285 UCB_CS285 Public

    My implementation of UCB cs285 deepRL homework

    Jupyter Notebook 76 7