We Scraped 1,407 Technical Job Listings from 6 LLM Labs — Here's Where R&D Is Actually Heading
Where is the current wave of LLM labs actually placing their technical bets? Keynotes and blog posts stay at the narrative layer. A more grounded signal is hiring JDs — the directions a company is willing to open headcount for are the directions it's genuinely willing to pay to bet on.
We pulled the public technical JDs from six companies — OpenAI, Anthropic, Zhipu AI, Moonshot AI, ByteDance, and DeepSeek — and after excluding non-technical roles (business, sales, legal, finance, HR, admin) and deduplicating, we were left with 1,407 technical roles, distributed as follows:
| Company | Technical JDs |
|---|---|
| ByteDance | 779 |
| OpenAI | 310 |
| Anthropic | 203 |
| Zhipu AI | 63 |
| DeepSeek | 27 |
| Moonshot AI | 25 |
Read these carefully: core model research is still the largest bucket, but Agent/Coding has already graduated from the "application layer" into its own track — the most clearly defined new priority of 2026.

R&D focus: core models still dominate, but Agent has spun out as its own track
Within those 1,407 technical roles we did a second pass, classifying by technical direction — keeping only core model / multimodal / Agent / Infra research / data & evals / safety / alignment directions, and excluding general backend / frontend / client engineering. That left 1,023 R&D roles. Broken down by direction, the structure looks like this:
| Direction | Total | % of R&D |
|---|---|---|
| Core model research / pretraining | 301 | 29.4% |
| Multimodal (vision / audio / video / embodied) | 246 | 24.0% |
| Agent / Coding | 183 | 17.9% |
| Training / inference Infra | 147 | 14.4% |
| Data / evaluation | 62 | 6.1% |
| Safety engineering | 62 | 6.1% |
| Post-training / alignment / RL | 22 | 2.2% |

A closer look at what this shows.
Core model research: the headcount gap is about org structure, not effort
Of the 301 core model research roles, ByteDance alone accounts for 235 (Seed, Doubao, TikTok Search, and Ads are all hiring large-model algorithm experts / algorithm leads), Zhipu AI has 14, and OpenAI and Anthropic each have 22–23. Read these numbers carefully. ByteDance's 235 doesn't mean it's "investing ten times the effort" — it means ByteDance splits model-algorithm roles very finely by business line (Lark, CapCut, Douyin, TikTok, Volcano Engine, Ads), with each BU hiring its own set. OpenAI and Anthropic concentrate the same kind of work inside their research orgs, where a single role covers a wider surface.
So the headcount distribution by country mainly reflects "org structure + business-line granularity," not "who is more willing to invest people in foundation models." The genuinely readable signal is elsewhere: Chinese labs are willing to open model-algorithm roles inside every BU, which means large-model capability is being staffed as shared infrastructure across business lines. OpenAI and Anthropic concentrate research inside their research orgs — closer to a "one team exporting capability to the outside" organizational shape.
Multimodal: ByteDance bets on both content generation and embodied; OpenAI bets only on embodied
Of the 246 multimodal roles, ByteDance alone contributes 217, spread across the "AI content generation" product line covering vision, speech, and video generation (mapped to Doubao, Jimeng, CapCut). OpenAI has fewer multimodal roles (around 17), and Anthropic barely touches the area (only Visual Knowledge Work and Audio, two research roles).
The real divergence is in embodied intelligence / robotics:
| Company | Embodied-related roles | Coverage |
|---|---|---|
| ByteDance | 23 (PICO / embodied) | VLA (vision-language-action) models, motion control, 3D simulation, XR vision, manipulation algorithms — staffed as product lines |
| OpenAI | 12 (dedicated Robotics) | Hardware (Electrical, Firmware), inference, safety, data collection — rebuilding a complete robotics team |
| Zhipu AI | 3 (X-lab) | Robot navigation / embodied algorithms |
ByteDance bets on both "content generation" and "embodied" as product directions; OpenAI bets only on "embodied robotics" research; Anthropic essentially doesn't touch embodied intelligence.
Agent / Coding: the most clearly defined new priority of 2026, spun out as its own track

Agent and Coding account for nearly a fifth of the 183 R&D roles, and the shape is strikingly consistent across the three leaders — invested in as "product lines," not "feature modules."
Agent is being run as an independent "model form" for post-training. OpenAI has carved "Agent Post-Training" into 8 independent directions: API & Power Users, Artifacts, Computer Use, Connectors, Context, Frontier Evals and Environments, Personality, Research. This slicing is almost identical to how, in the GPT era, chat models were split by use case. Inside OpenAI, agent is treated as a distinct model form — not an attached feature of some product.
"Harness" is becoming a role category in China too. Moonshot AI lists Harness Development Engineer and Harness Research Engineer; ByteDance has attached a dozen "AI Agent Harness Optimization Specialist" roles under TRAE and Jimeng; DeepSeek has staffed Agent Harness Development Engineer, Researcher, and Product Manager as a three-role set. Anthropic popularized the concept of the "agent harness" (the agent runtime framework) over the past two years, and Chinese labs are now systematically building teams around it.
Coding Agent has become its own track, with each lab staffing dedicated R&D, product, and solutions roles:
| Company | Coding product | Headcount |
|---|---|---|
| OpenAI | Codex | 20 Codex roles + 1 Coding Agents security role |
| ByteDance | TRAE | 18 TRAE-dedicated roles (36 coding-related in total) |
| Anthropic | Claude Code | 6 (including a dedicated Model Performance Engineer) |
| Zhipu AI / Moonshot AI | GLM Coding Agent / Kimi Code | Each listing roles independently |
Computer Use: agents move from "calling APIs" to "operating real environments"

"Computer Use" appears as an independent direction in the hiring of both OpenAI and Anthropic. OpenAI has Agent Post-Training, Computer Use Research and Software Engineer, Computer Use & Frontier Interfaces; Anthropic has Product Engineer, Computer Use and Research Engineer, Computer Use. Both leaders treat it as a direction worth staffing independently — a signature signal that agents are moving from "calling APIs" to "operating real environments."
Two new job types Agent has given rise to: security, evals

Once agents can run tasks, call tools, and operate environments on their own, "model alignment" is no longer enough — and two job types that didn't previously exist have emerged.
Agent security. Anthropic has built Safeguards into a standalone 13-person system: Safeguards Labs, Evals, Infrastructure, Policy, Review Tooling, Rare Harms. OpenAI's counterparts are Agent Security, Offensive Security Engineer Agent Products, Security Researcher Agentic AI Threats, Agentic Risk Analyst, plus a dedicated Security Preparedness Lead, Coding Agents. What they're watching is no longer "is the model's output aligned," but a new question: "is the agent's behavior safe."
Evaluation infrastructure. In the model era, evaluation meant running benchmarks. In the agent era, evaluation means building environments, running trajectories, and judging outcomes — an order of magnitude more complex. ByteDance alone has 28 evaluation roles, split finely by business line (Douyin, CapCut, Lark, Doubao, Seed, Volcano Ark, medical — each its own line), including a dedicated "Large Model and Agent Evaluation Infrastructure Algorithm Engineer." OpenAI and Anthropic each have 4–7 eval roles, and Moonshot AI and Zhipu AI both list independent Eval Engineers. Evaluation has spun out from research into a separately funded direction.
Post-training / alignment / RL: rarely an independent role, but not under-hired
By title keyword, only 22 roles have Alignment / RLHF directly in the title. That number is low because post-training / alignment work isn't broken out into independent direction-level roles in the JDs — it's folded into other buckets: OpenAI's "Agent Post-Training" series does the post-training / RL of the agent era, but by title it lands in the Agent bucket; ByteDance's "large-model application algorithm" roles contain a large amount of SFT (supervised fine-tuning) / RL work but are filed under core model research; and these roles are generally listed under a generic "Researcher" title rather than a direction title. In other words, post-training / alignment is something every lab does — it just hasn't been organized into an independent direction sequence.
Infra: the inference-first turn is now unmistakable
Training / inference Infra accounts for 147 R&D roles. Slice the direction words more finely:
| Company | Inference-related | Training-related |
|---|---|---|
| OpenAI | 9 | 13 |
| Anthropic | 14 | 4 |
| ByteDance | 45 | 36 |
| Zhipu AI | 6 | 9 |
| Moonshot AI | 1 | 1 |

Anthropic's inference headcount is 3.5× its training headcount, and OpenAI is also clearly shifting its center of gravity toward inference (including data-center roles like Stargate, Rack Infrastructure, Site Operations). The driver: once model capability converges and the price war starts, every lab has to drive inference cost down. ByteDance's 45 inference roles correspond to the scale at which Volcano Engine's Ark MaaS (model-as-a-service) serves tokens to the outside.
"Cheap inference" has already collided with the hardware layer. Look at the specific roles: Zhipu AI's inference Infra engineer listing explicitly bundles "quantization algorithm research + inference framework optimization + GPU optimization" into one; Anthropic is hiring a TPU Kernel Engineer and a Performance Engineer, GPU; OpenAI is working on AMD GPU Enablement (de-NVIDIA-fying inference), Kernel Performance & AI Tooling, and Hardware/Software CoDesign; ByteDance's Seed Infra is hiring LLM/VLM Inference Optimization researchers. Tweaking the framework layer is no longer enough — you have to push down into the kernel and quantization algorithms, and even start using AMD / TPU / in-house ASICs to share out NVIDIA's compute cost.
In closing
Compress 1,407 technical JDs into one sentence: the industry's technical center of gravity is moving from "train bigger models" to "train agents that are better at doing things."
Core model research is still the biggest bucket (29%), multimodal is second (24%), but Agent/Coding (18%) is the most clearly defined new priority of 2026 — OpenAI has already carved agent post-training into 8 independent directions, and Computer Use, Agent security, and evaluation infrastructure have all become newly independent job types. On the Infra side, the inference-first turn is written into headcount, and optimization has already collided with the kernel and quantization layer.
This data has its limits, of course: the real share of post-training / RL is understated by title-naming conventions; headcount isn't the same as investment intensity — a sizable portion of ByteDance's 235 core model roles are duplicated allocations across business lines. To judge "who is more aggressive," counting JDs alone isn't enough; you also have to look at each lab's fundraising cadence and compute reserves.
But one thing is clear: the structural shift in this round of 2026 hiring says, more honestly than any PR release ever could, what LLM labs intend to do next. Agent is no longer a vision in a slide deck — it's written into headcount, written into independent hiring directions. The roles now being carved out as standalone job types — Computer Use, Agent security, evaluation infrastructure — are very likely the products we'll be using every day next year.
Seen that way, these 1,407 JDs aren't just a hiring snapshot — they're more like a roadmap. The year ahead is well worth watching.
