You will contribute to the development of high-performance kernels, optimize machine learning models, and build infrastructure for hardware-software integration. The role involves working across teams to manage memory, task scheduling, and distributed communication for AI systems.
Requirements overview
Candidates must be currently pursuing a BS, MS, or PhD in a technical field with experience in parallel processing, machine learning, or distributed systems. Proficiency in C/C++, Python, and ML frameworks like PyTorch is required, along with an understanding of optimization techniques.
Tenstorrent is a next-generation computing company that builds computers for AI.
Headquartered in the U.S. with offices in Austin, Texas, and Silicon Valley, and global offices in Toronto, Belgrade, Seoul, Tokyo, and Bangalore, Tenstorrent brings together experts in the field of computer architecture, ASIC design, RISC-V technology, advanced systems, and neural network compilers. Tenstorrent is backed by Eclipse Ventures and Real Ventures, Archerman Capital, Samsung Catalyst Fund, and Hyundai Motor Group among others.
Join us: www.tenstorrent.com/careers.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.
Overview: This is the team that makes "it runs on Tenstorrent hardware" actually true for real models, at real scale, and builds the infrastructure that keeps the rest of engineering moving fast.
This role is on-site based out of Austin, TX or Santa Clara, CA.
What you might work on: This posting spans multiple teams within Kernels, Models, Inference, Scaleout & our Runtime teams. One application, one recruiter screen, then we match you to the specific team and location that fits best. Teams include:
Kernels: Develops high performance kernels on Tenstorrent hardware
Models: Optimizes ML models (LLMs, vision models, video and image generation, and other architectures) for our hardware
Inference Server: development and serving-side optimization
Runtime: Builds the software engine that manages memory, task scheduling, and code execution on hardware while an application is actively running.
Scale Out: communication and coordination between devices in a distributed AI system.
Compiler & Infra: Create tools that optimize AI models into high-performance programs on Tenstorrent hardware, covering memory planning, profiling, debug, and emulation.
Who You Are
Currently pursuing a BS, MS, or PhD in Comp Sci, Comp Eng, Physics/Math or a related field.
Coursework or projects any of the following: parallel processing, machine learning or distributed systems.
A high comfort level as an end user of AI and agentic flows.
What We Need
Experience with model quantization, kernel fusion, or other optimization techniques; or low level high performance programming.
Hands on experience working with Pytorch, experimenting with and deploying models, for inference or training.
Experience with C/C++ or kernel development in languages like CUDA, or Python and at least one ML framework (PyTorch, TensorFlow, JAX).
Knowledge of RTL, HDL is nice to have on certain teams.
What You Will Learn
How models get adapted to run well on specialized hardware
How to write high performance code that can get the most out of the underlying hardware
Production ML serving infrastructure at scale
How infrastructure choices affect the velocity of an entire engineering org
USA Hiring Timelines
This internship opportunity is available throughout our 3 terms with the following corresponding recruitment cycles:
Winter Term: Jan–Apr work term, Sept–Dec recruit.
Summer Term: May–Aug work term, Oct–Apr recruit.
Fall Term: Sept–Dec work term, Jan–Aug recruit.
Please note these timelines are for reference only. Actual timelines may vary.
Compensation for all interns at Tenstorrent ranges from $50/hr - $70/hr including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made.
Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.
This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.