folks i promise i will start writing blogs more often. it's been 3 weeks since the last update. there have been some new things i want to share, and one of them is yc startup school afterthoughts — so here we go.
it wasn't in my initial plan to attend because i was placed on a waitlist a week prior to the event date. coincidentally, on the first day of startup school, i was in town meeting an old friend for lunch. while i was walking towards the ramen restaurant i receive an email notification that yc has accepted me into startup school. i was very surprised because i didn't expect this to happen on the day of. oh well, it's a good opportunity and i am taking it, so i booked a lyft to chase center right after.
biggest takeaway: physical intelligence
i went half for the talks, half for the merch, but chelsea finn's talk on physical ai is the thing i actually walked out thinking about. she's a stanford professor and one of the people behind physical intelligence, and it mostly just solidified something i'd been leaning toward anyway — i want to work on hardware.
her framing was that physical ai is roughly where computer vision was before pretraining. until ~2023, robotics basically meant collecting a bespoke dataset for your one task and training from scratch every time, like re-collecting imagenet on every project. that's the bottleneck. no shared foundation, no transfer.
the way out is the same move language and vision already made: generalist pretrained models, in this case vision-language-action models (VLAs). the reliability problem was the interesting part. you can get a robot to make espresso around 90% of the time, and the last few nines are the actual work. usually a human iterates on the dataset, more data, cleaner labels, edge cases. it works, though as she put it, the person eventually gets tired. the more interesting version is the system iterating on itself, finding where it needs more supervision automatically. that's roughly their π*0.6 work: a VLA that learns from its own experience, with a human only stepping in to correct. she framed that automatic loop as maybe the only realistic path to 99.9%.
i want to read more about this. i caught the shape of it in the talk and want to understand the model architectures properly, the VLAs, the value function, how the memory pieces fit together. more to come once i dive deep into it.
after parties rant
i want to start off by saying networking is not my forte, so i actually had some doubts about going to these after parties. but, a guy gotta do what he gotta do. i went to the parties hosted by lemma, aws (attempted, but line was genuinely long enough to the point that i gave up), and google deepmind. the venues were all super duper nice with pretty decent food and drinks. also met some new people and caught up with some friends! overall would rate a 8/10.
wrap up
that was pretty much startup school for me. i showed up on a waitlist acceptance i got a few hours before, sat in a room full of cracked builders, and left with one talk i can't stop thinking about plus a tote bag i definitely didn't need. fair trade.
the physical ai stuff is the thing i'm actually carrying out of it. everything else was fun, but that's the part pointing me somewhere. and yeah, i said i'd write more often. consider this me trying.
— yihan
