A number of companies are and have tried to build this. I myself did once in 2018. Alyc.ai if you care to look. We had voice, vision, llm, domain specific training (called fine tuning today I suppose), emotion detection, gesture detection, pose estimation, action recognition (ie. guy is drinking a beer), multiple stereoscopic cameras, microphone mesh, infrared camera (used for depth and night vision), and used nvidia chipsets (jetson) to run models at the edge. Other models ran in the cloud. Our LLM was about 2B parameters trained on about 800MB of Japanese data. Also our interface was a sophisticated peppers ghost (Today I would use lightfield displays) but it was fun. We didn't have RLHF feedback loops so Alyc was a bit complex. Still people loved her. Covid killed this project.
Comments
A number of companies are and have tried to build this. I myself did once in 2018. Alyc.ai if you care to look. We had voice, vision, llm, domain specific training (called fine tuning today I suppose), emotion detection, gesture detection, pose estimation, action recognition (ie. guy is drinking a beer), multiple stereoscopic cameras, microphone mesh, infrared camera (used for depth and night vision), and used nvidia chipsets (jetson) to run models at the edge. Other models ran in the cloud. Our LLM was about 2B parameters trained on about 800MB of Japanese data. Also our interface was a sophisticated peppers ghost (Today I would use lightfield displays) but it was fun. We didn't have RLHF feedback loops so Alyc was a bit complex. Still people loved her. Covid killed this project.