HirePortal

AI Infrastructure Engineer / Architect (Infra & GPU Computing)

  • VAST
  • Shanghai, Beijing, China
  • CNY 600,000 – CNY 1,200,000

AI Infrastructure EngineerAI基础设施工程师/架构师(Infra&GPU Computing)

Shanghai, BeijingExperiencedFull-time

Responsibilities

Job Description What We Are Building 1.TripoAI is building a world-leading generative AI platform for 3D content creation. Behind our product, thousands of high-performance GPUs run complex inference computations around the clock. 2.For the Infrastructure team, our mission goes far beyond merely keeping systems operational — we aim to redefine cloud computing architecture for the AI era. 3.We are looking for a technical partner who can manage large-scale GPU clusters, maintain acute cost awareness, and pursue ultimate performance to build Tripo’s global computing infrastructure. If you are tired of maintaining legacy business systems at large corporations and eager to design an AI-native high-performance computing cluster from scratch, this is your battlefield. Challenges & Core Responsibilities 1.You will serve as the primary owner of Tripo’s computing resource efficiency, directly shaping the company’s technical cost structure and the ceiling of user experience. 2.Build AI computing scheduling systems: Design and optimize GPU task scheduling platforms based on Kubernetes and Ray. Tackle tough technical challenges in LLM inference such as video memory isolation, cold start optimization, and queuing policies to maximize GPU utilization. 3.Global high-availability SRE architecture: Design multi-region traffic scheduling and disaster recovery solutions to guarantee 99.9% SLA and ultra-low latency for global users accessing Tripo APIs. 4.Extreme cost optimization (FinOps): You will act as both engineer and cost specialist. Leverage mixed deployment of Spot instances, HPA/VPA auto-scaling, model quantization deployment and other approaches to slash inference costs while sustaining stable performance. 5.Develop AI developer experience (DevEx) platforms: Build automated CI/CD pipelines and model training/evaluation platforms for algorithm and backend teams to shorten the entire workflow from code commit to production launch. 【我们在做什么】 1、TripoAI正在构建全球领先的3D生成式AI平台。我们的产品背后,是数以千计的高性能GPU在日夜不停地进行复杂的推理计算。 2、对于Infra团队,我们的使命不仅是“系统不挂”,而是重新定义AI时代的云计算架构。 3、我们需要一位能驾驭大规模GPU集群、对成本极其敏感、对性能有极致追求的技术合伙人,来构建Tripo的全球算力底座。 如果你厌倦了在大厂维护陈旧的业务系统,渴望亲手设计一套面向AI Native的高性能计算集群,这里是你的战场。 【你将面临的挑战与职责】 1、你将是Tripo算力效能的第一负责人,直接决定公司的技术成本结构和用户体验上限。 2、构建AI算力调度系统:基于K8s/Ray等技术,设计并优化GPU任务调度平台。解决大模型推理中的显存隔离、冷启动优化、排队策略等硬核难题,将GPU利用率提升至极致。 3、全球化高可用架构(SRE):设计跨区域(Multi-Region)的流量调度与容灾方案,确保全球用户访问TripoAPI时的高可用性(SLA 99.9%)和低延迟。 4、极致的成本优化(FinOps):你不仅是工程师,也是精算师。通过Spot实例混部、自动扩缩容(HPA/VPA)、模型量化部署等手段,在保证性能的前提下,大幅降低推理成本。 5、打造AI研发效能平台(DevEx):为算法和后端团队构建自动化的CI/CD流水线、模型训练/评测平台,让代码从提交到上线的路径最短化。

Qualifications

Job Requirements Who We Are Looking For We do not value rigid rote knowledge; we prioritize complex problem-solving capabilities and an entrepreneurial ownership mindset. Hard Core Competencies 1.Solid technical foundation: Proficient in at least one language among Go, Python and Java; possess in-depth understanding of distributed systems, microservice architecture, and databases (SQL/NoSQL/VectorDB). 2.Excellent architecture taste: Hold strong code taste, prioritize code simplicity and maintainability, and write elegant code recognized by the whole engineering team. 3.End-to-end delivery capability: Hands-on experience building high-concurrency, high-availability systems; familiar with cloud-native stacks including Kubernetes, Docker, AWS/GCP; able to independently oversee full-cycle work from architecture design to production release. Entrepreneurial Traits (Preferred Qualifications) 1.High self-motivation: Require minimal supervision; proactively identify issues and design corresponding solutions. You focus not only on closing tickets but also on delivering tangible business value. 2.Comfort with uncertainty: Able to rapidly adjust technical decisions amid fast-paced startup environments, striking a balance between perfect architecture and rapid iteration. 3.AI enthusiast: More than just an AI end-user — you conduct in-depth research on AI and aspire to reshape software development workflows with AI technology. What We Offer 1.Core team influence: Flat organizational structure. Every technical decision you make will directly steer product development and business metrics. There are no cumbersome reporting procedures, only fast feedback loops. 2.Steep growth trajectory: Tackle cutting-edge global 3D + AI technical challenges with no established industry solutions, and grow alongside a team of elite technical specialists. 3.AI-first geek culture: We reject meaningless internal competition and advocate efficient workflows powered by standardized SOPs and AI assistants. We encourage experimentation with new technologies and pursue refined, high-quality solutions rather than merely meeting functional requirements. 4.Competitive compensation package: Besides base salary, we offer equity tied to the company’s growth for long-term partners who wish to grow with us. 【我们在寻找这样的你】 我们不看重死记硬背的八股文,我们看重解决复杂问题的能力和创业者的Owner意识。 硬核素质: 1、技术底座扎实:精通Go/Python/Java其中至少一门语言,对分布式系统、微服务架构、数据库(SQL/NoSQL/VectorDB)有深刻理解。 2、架构设计品位:具备优秀的代码品位(CodeTaste),追求代码的简洁性与可维护性,能写出让团队赞叹的优雅代码。 3、工程落地能力:有高并发、高可用系统的实战经验,熟悉云原生技术栈(K8s,Docker,AWS/GCP),能独立Cover从设计到上线的全流程。 创业者基因(加分项): 1、极强的自驱力(Self-driven):不需要被管理,能主动发现问题并定义解决方案。你关注的不只是Ticket是否关闭,而是业务价值是否交付。 2、拥抱不确定性:在快速变化的创业环境中,能迅速调整技术决策,在“完美架构”与“快速迭代”之间找到平衡。 3、AI信仰者:不仅是AI的使用者,更是AI的深度研究者,渴望用AI重塑软件开发流程。 【你将获得什么】 1、核心团队的影响力:扁平化管理,你的每一个技术决策都能直接影响产品走向和业务数据,没有繁琐的汇报流程,只有极速的反馈闭环。 2、陡峭的成长曲线:直接面对全球最前沿的3D+AI技术场景,解决行业内没有标准答案的难题,与一群技术极客共同进化。 3、AI驱动的极客氛围:我们拒绝低效内卷,推崇“SOP+AI辅助”的高效工作流。在这里,我们鼓励尝试新技术,致力于把事情做漂亮,而不仅仅是做完。 4、有竞争力的回报:除了薪资,我们提供与公司成长绑定的期权,寻找愿意长期同行的伙伴。

Apply

Skills

  • Kubernetes
  • GPU Computing
  • Ray
  • SRE
  • LLM inference optimization
  • Cloud Architecture
  • Cost Optimization

Related jobs

VASTApply for this job