Pipeline run
48c27ce6-5157-4239-ae3b-7d5979a3f662
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
Roku is looking for a Senior Engineer to join their backend and data team, focusing on designing and optimizing distributed data pipelines and real-time processing systems. The role requires extensive experience in Java, distributed systems, and big data technologies, along with a strong background in building scalable streaming solutions. Candidates should also demonstrate proficiency in AI tools and automation, showcasing their ability to drive innovation and improve processes. Roku operates in the streaming industry, emphasizing a culture of collaboration and problem-solving to enhance user experiences.
Job description
What does the team work on? Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television – including TV advertising – around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers' experiences. What is the role? We are seeking a highly experienced and skilled Senior Engineer to join our backend and data team. This role is crucial for designing, building, and optimizing distributed data pipelines, real-time data processing systems, and backend solutions that effectively handle large-scale data. The ideal candidate will have deep expertise in Java, distributed systems, and big data technologies, and a passion for solving complex problems and delivering robust solutions. We’re always in “build mode” because we’re a company of data-focused builders. Every day, you’ll look at what exists and find ways to make it better and help drive innovation. How will I use AI at Roku? At Roku, we don’t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We’re looking for curious, adaptable builders who can show how they’ve used AI or automation to move faster, raise the bar, and scale their impact. We value your AI skills if have built fluency across the agentic engineering toolchain — coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you. What are the responsibilities of the role? • Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale • Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence • Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations • Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions • Remain significantly hands-on with critical features and architectural components • Drive technical planning, prioritization, and execution aligned with business objectives • Act as a key technical partner to the Engineering Manager in driving team success and technical decisions • Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices • Lead efforts in monitoring, observability, performance optimization, and production reliability at scale What experience would help someone be successful in this role at Roku? • 8+ years of software engineering experience with significant time in senior engineering roles • Proven expertise in building scalable, distributed, and streaming solutions in production environments • Deep experience with content management systems, media processing, or publishing platforms • Expert-level proficiency in Java or Scala required; Python experience is a strong plus • Strong expertise in distributed systems architecture, microservices, and event-driven architectures • Deep understanding of streaming technologies (Kafka, Redpanda, or similar • Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale • Track record of leading complex technical projects from conception to production in high-scale environments • Excellent communication skills and ability to influence technical decisions across teams and organizations • Experience with cloud platforms (AWS, GCP, or Azure) at enterprise scale • Extensive experience with containerization and orchestration (Docker, Kubernetes) • Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems #LI-AB3 What's Roku's approach to hybrid working? Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days' a week attendance. What are some of the benefits? Roku is committed to offering a diverse range of benefits as part of our compensation package to support our employees and their families. Our comprehensive benefits include global access to mental health and financial wellness support and resources. Local benefits include statutory and voluntary benefits which may include healthcare (medical, dental, and vision), life, accident, disability, commuter, and retirement options (401(k)/pension). Employees are supported in taking time off, in accordance with local leave policies and other personal needs to support their evolving work and life needs. It's important to note that not every benefit is available in all locations or for every role. For details specific to your location, please consult with your recruiter. Accommodations Roku welcomes applicants of all backgrounds and provides reasonable accommodations and adjustments in accordance with applicable law. If you require reasonable accommodation at any point in the hiring process, please direct your inquiries to EmployeeRelations@Roku.com. What should I know about Roku's culture? Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company's success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We're independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you'll be part of a company that's changing how the world watches TV. We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn't real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002. To learn more about Roku, our global footprint, and how we've grown, visit https://www.weareroku.com/factsheet. By providing your information, you acknowledge that you want Roku to contact you about job roles, that you have read Roku's Applicant Privacy Notice, and understand that Roku will use your information as described in that notice. If you do not wish to receive any communications from Roku regarding this role or similar roles in the future, you may unsubscribe at any time by emailing WorkforcePrivacy@Roku.com
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
This skill is in Secondary as the JD says "Experience with search technologies (Elasticsearch, Solr, etc)" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "While we are well-positioned to help shape the future of television – including TV advertising – around the world, continued success relies on its investment in our machine learning capabilities." - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "Strong expertise in distributed systems architecture, microservices, and event-driven architectures" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "vector databases (e.g., Milvus)" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "MCP servers, custom skills, or agent frameworks" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "Advanced knowledge of databases (SQL/NoSQL)" - it isn't in this role's skill catalog yet, so it's tracked for review.
This skill is in Secondary as the JD says "Python experience is a strong plus" - the JD marks it as "a plus" rather than required.
This skill is in Secondary as the JD says "• Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
This skill is in Secondary as the JD says "Expert-level proficiency in Java or Scala required" - the JD marks it as "a plus" rather than required.
This skill is in Secondary as the JD says "• Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Slowly Changing Dimensions — Data Modeling & Warehouse Design (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": {
"source_marker": {
"first_5_words": "Roku pioneered streaming to the",
"last_5_words": "shape customers\u0027 experiences."
},
"text": "Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers\u0027 experiences.",
"word_count": 84
},
"ai_kras": [
"We value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks.",
"You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you."
],
"certifications": [],
"client_details": null,
"company_name": "Roku",
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [
"Streaming",
"Television"
],
"domain": "Media \u0026 Entertainment"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 8,
"raw": "8+ years of software engineering experience with significant time in senior engineering roles"
},
"job_locations": [
{
"aliases": [],
"city": null,
"country": null,
"state": null,
"work_mode": "hybrid"
}
],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Senior Engineer",
"role_aliases": [
{
"name": "Senior Software Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Backend Engineer",
"reasoning": "focus on backend solutions and data pipelines",
"relation": "synonym"
},
{
"name": "Data Engineer",
"reasoning": "role involves building and optimizing data pipelines",
"relation": "synonym"
}
],
"role_archetype": "Engineering",
"roles_and_responsibilities": [
{
"bullet_count": 9,
"heading": "Responsibilities of the role",
"heading_was_present": true,
"source_marker": {
"first_5_words": "\u2022 Define architecture and technical",
"last_5_words": "monitoring, observability, performance optimization"
},
"text": "\u2022 Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale\n\u2022 Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence\n\u2022 Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations\n\u2022 Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions\n\u2022 Remain significantly hands-on with critical features and architectural components\n\u2022 Drive technical planning, prioritization, and execution aligned with business objectives\n\u2022 Act as a key technical partner to the Engineering Manager in driving team success and technical decisions\n\u2022 Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices\n\u2022 Lead efforts in monitoring, observability, performance optimization, and production reliability at scale",
"word_count": 174
},
{
"bullet_count": 0,
"heading": "How will I use AI at Roku?",
"heading_was_present": true,
"source_marker": {
"first_5_words": "At Roku, we don\u2019t just",
"last_5_words": "with an agent helping you."
},
"text": "At Roku, we don\u2019t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We\u2019re looking for curious, adaptable builders who can show how they\u2019ve used AI or automation to move faster, raise the bar, and scale their impact.\nWe value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you.",
"word_count": 114
}
],
"urls": [
{
"type": "website",
"url": "https://www.weareroku.com/factsheet"
}
]
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
},
{
"dimension": "Data Modeling \u0026 Warehouse Design",
"reason": "In the role canon\u0027s Data Modeling \u0026 Warehouse Design - adding it narrows the candidate pool.",
"skill": "Slowly Changing Dimensions",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "3f686ed1-fac0-5707-97b5-c06d4b5e8e4c",
"skill_name": "Java"
},
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;",
"skill_id": "a2275eea-20e3-59dd-b201-af76b7ae706f",
"skill_name": "Apache Kafka"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
"skill_id": "78c5a4b9-2ab5-5279-a688-46f4c3767c3f",
"skill_name": "AWS"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
"skill_id": "de9d33d1-f608-5b2a-97b7-7398054f33cc",
"skill_name": "Google Cloud Platform"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
"skill_id": "060b1d39-d4d8-581d-9d96-65cd45d7cea4",
"skill_name": "Microsoft Azure"
},
{
"dimension": {
"display_name": "Containers \u0026 Orchestration",
"slug": "containers-orchestration"
},
"is_primary": false,
"layer": "L2",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.",
"skill_id": "37ac0455-c683-5062-9393-96c9284f5ed1",
"skill_name": "Docker"
},
{
"dimension": {
"display_name": "Containers \u0026 Orchestration",
"slug": "containers-orchestration"
},
"is_primary": false,
"layer": "L2",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.",
"skill_id": "57e0c6c2-06e3-5d5a-a0cd-249da956a7e3",
"skill_name": "Kubernetes"
},
{
"dimension": {
"display_name": "Observability \u0026 Monitoring",
"slug": "observability-monitoring"
},
"is_primary": false,
"layer": "L2",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_5 assigns run-health monitoring of owned pipelines: run metrics, log triage, and alerting on failures and SLA misses.",
"skill_id": "48406a13-f524-5e91-af33-a3ef313fe54d",
"skill_name": "Elasticsearch"
},
{
"dimension": {
"display_name": "Message Brokers",
"slug": "message-brokers"
},
"is_primary": false,
"layer": "L2",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Native by registry design: the Java catalog tiered this dimension borrowed(Data Engineer), anticipating this role as its owner.",
"skill_id": "456197d8-7e5d-5f86-9069-411b8eafe2c2",
"skill_name": "Event-Driven Architecture"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "dbe274b4-5b86-599d-bac5-1fc85691053a",
"skill_name": "Big Data"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review.",
"skill_id": "08775033-5e69-55de-b80d-518b6143050b",
"skill_name": "Redpanda"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review.",
"skill_id": "f683d14a-c143-528d-92d6-f4bca348cb58",
"skill_name": "Claude Code"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review.",
"skill_id": "82273e02-f063-501f-ac18-777e95151380",
"skill_name": "Cursor"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": "Roku",
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 8,
"raw": "8+ years of software engineering experience with significant time in senior engineering roles"
},
"industryDomain": "Media \u0026 Entertainment",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 8
},
"location": [
"Hybrid"
],
"noticePeriod": null
},
"locations": [
"Hybrid"
],
"noticePeriod": null,
"openToRelocate": false,
"role": "Senior Engineer",
"roleSynonyms": [
"Senior Software Engineer",
"Backend Engineer",
"Data Engineer"
]
},
"jd_summary": "Roku is looking for a Senior Engineer to join their backend and data team, focusing on designing and optimizing distributed data pipelines and real-time processing systems. The role requires extensive experience in Java, distributed systems, and big data technologies, along with a strong background in building scalable streaming solutions. Candidates should also demonstrate proficiency in AI tools and automation, showcasing their ability to drive innovation and improve processes. Roku operates in the streaming industry, emphasizing a culture of collaboration and problem-solving to enhance user experiences.",
"kras": [
{
"cluster": "Data Pipeline \u0026 Warehouse Ownership",
"evidence": {
"quote": "Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations",
"similarity": 0.5176
},
"kra_id": "2f63ae9b-f6f9-58cd-80a6-649e080ef381",
"role_name": "Data Engineer",
"role_phrasing": "Deliver the pipelines: batch ETL/ELT in Spark/Python/SQL, orchestrated as owned DAGs, landing conformed datasets in lake and warehouse on schedule.",
"role_slug": "data-engineer",
"salience": 1.0,
"source": "jd",
"weight": 1.0
},
{
"cluster": "Data Architecture \u0026 Modeling",
"evidence": {
"quote": "Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations",
"similarity": 0.558
},
"kra_id": "f815ce00-9fa0-5baf-8bc9-b74889f4bc2c",
"role_name": "Data Engineer",
"role_phrasing": "Design where and how data lives: lakehouse layout, open table formats, partitioning, and dimensionally modeled target structures with deliberate grain and history handling.",
"role_slug": "data-engineer",
"salience": 0.85,
"source": "jd",
"weight": 0.85
},
{
"cluster": "Data Ingestion \u0026 Integration",
"kra_id": "50c92273-549b-5795-9a4a-0ac738d9c5f7",
"role_name": "Data Engineer",
"role_phrasing": "Own how data enters the platform: connector-based EL, change-data-capture, API extraction, and streaming ingest basics \u2014 incremental, deduplicated, and schema-aware.",
"role_slug": "data-engineer",
"salience": 0.8,
"source": "graph",
"weight": 0.8
},
{
"cluster": "Data Quality \u0026 Reliability Assurance",
"kra_id": "0b59e7cd-a4ad-5a60-b75b-3b88ea833f36",
"role_name": "Data Engineer",
"role_phrasing": "Make trustworthiness testable: expectation suites and contract checks in pipelines, freshness/volume/anomaly monitors on published datasets, and quarantine paths for bad data.",
"role_slug": "data-engineer",
"salience": 0.75,
"source": "graph",
"weight": 0.75
},
{
"cluster": "Production Support \u0026 Incident Resolution",
"evidence": {
"quote": "Lead efforts in monitoring, observability, performance optimization, and production reliability at scale",
"similarity": 0.4748
},
"kra_id": "f24d5f4f-042b-5b0b-b3ed-9b7dbca5c21d",
"role_name": "Data Engineer",
"role_phrasing": "Keep owned pipelines healthy in production: triage failed runs and late data fast, execute clean backfills, and close incidents with root-cause fixes.",
"role_slug": "data-engineer",
"salience": 0.7,
"source": "jd",
"weight": 0.7
},
{
"cluster": "Automated Testing \u0026 Code Quality",
"evidence": {
"quote": "Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence",
"similarity": 0.4907
},
"kra_id": "7fc0a198-8285-569f-8ea5-1a97793ebb6a",
"role_name": "Data Engineer",
"role_phrasing": "Treat pipeline code as production software: pytest coverage on transforms and DAG logic, integration tests against sample data, and substantive review of pipeline PRs.",
"role_slug": "data-engineer",
"salience": 0.65,
"source": "jd",
"weight": 0.65
},
{
"cluster": "Data Platform Cost \u0026 Performance Optimization",
"kra_id": "03f9a4ac-da98-556e-b928-750909b67955",
"role_name": "Data Engineer",
"role_phrasing": "Own the efficiency of owned workloads: tune Spark jobs and warehouse loads, right-size compute, manage partitioning and file compaction, and hold storage/egress costs in check.",
"role_slug": "data-engineer",
"salience": 0.6,
"source": "graph",
"weight": 0.6
},
{
"cluster": "Data Security \u0026 Access Compliance",
"kra_id": "dfdfbcd0-ecbf-57dc-b9ad-ba5ecf9e13e7",
"role_name": "Data Engineer",
"role_phrasing": "Handle data like it matters: scoped IAM on owned buckets and warehouses, encryption at rest and in transit, secrets hygiene in orchestrators, and PII-aware pipeline design.",
"role_slug": "data-engineer",
"salience": 0.5,
"source": "graph",
"weight": 0.5
},
{
"cluster": "Codebase Health \u0026 Delivery Hygiene",
"evidence": {
"quote": "Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence",
"similarity": 0.4929
},
"kra_id": "05ddbc30-a16f-5add-a4cf-052dcc6b20ab",
"role_name": "Data Engineer",
"role_phrasing": "Keep pipeline codebases modern and shippable: refactor legacy jobs incrementally, keep engine/orchestrator versions current, and keep CI on pipeline repos fast and green.",
"role_slug": "data-engineer",
"salience": 0.45,
"source": "jd",
"weight": 0.45
},
{
"cluster": "Technical Documentation",
"kra_id": "d6a89be2-94b2-5ad1-93c0-236b130eafe3",
"role_name": "Data Engineer",
"role_phrasing": "Document owned datasets and pipelines so consumers self-serve: dataset descriptions and grain, pipeline runbooks, and short decision records.",
"role_slug": "data-engineer",
"salience": 0.2,
"source": "graph",
"weight": 0.2
}
],
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": {
"source_marker": {
"first_5_words": "Roku pioneered streaming to the",
"last_5_words": "shape customers\u0027 experiences."
},
"text": "Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers\u0027 experiences.",
"word_count": 84
},
"ai_kras": [
"We value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks.",
"You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you."
],
"certifications": [],
"client_details": null,
"company_name": "Roku",
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [
"Streaming",
"Television"
],
"domain": "Media \u0026 Entertainment"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 8,
"raw": "8+ years of software engineering experience with significant time in senior engineering roles"
},
"job_locations": [
{
"aliases": [],
"city": null,
"country": null,
"state": null,
"work_mode": "hybrid"
}
],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Senior Engineer",
"role_aliases": [
{
"name": "Senior Software Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Backend Engineer",
"reasoning": "focus on backend solutions and data pipelines",
"relation": "synonym"
},
{
"name": "Data Engineer",
"reasoning": "role involves building and optimizing data pipelines",
"relation": "synonym"
}
],
"role_archetype": "Engineering",
"roles_and_responsibilities": [
{
"bullet_count": 9,
"heading": "Responsibilities of the role",
"heading_was_present": true,
"source_marker": {
"first_5_words": "\u2022 Define architecture and technical",
"last_5_words": "monitoring, observability, performance optimization"
},
"text": "\u2022 Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale\n\u2022 Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence\n\u2022 Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations\n\u2022 Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions\n\u2022 Remain significantly hands-on with critical features and architectural components\n\u2022 Drive technical planning, prioritization, and execution aligned with business objectives\n\u2022 Act as a key technical partner to the Engineering Manager in driving team success and technical decisions\n\u2022 Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices\n\u2022 Lead efforts in monitoring, observability, performance optimization, and production reliability at scale",
"word_count": 174
},
{
"bullet_count": 0,
"heading": "How will I use AI at Roku?",
"heading_was_present": true,
"source_marker": {
"first_5_words": "At Roku, we don\u2019t just",
"last_5_words": "with an agent helping you."
},
"text": "At Roku, we don\u2019t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We\u2019re looking for curious, adaptable builders who can show how they\u2019ve used AI or automation to move faster, raise the bar, and scale their impact.\nWe value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you.",
"word_count": 114
}
],
"urls": [
{
"type": "website",
"url": "https://www.weareroku.com/factsheet"
}
]
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"election_confidence": 0.0,
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"role_coverage": [
{
"is_jd_role": true,
"name": "Data Engineer",
"pct": 0.37,
"role": "data-engineer",
"share": 100.0
}
],
"run_id": "jdv4-cf40a75f2bf0",
"secondary_meta": [
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Experience with search technologies (Elasticsearch, Solr, etc)\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "Experience with search technologies (Elasticsearch, Solr, etc)",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Apache Solr"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities.\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities.",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Machine Learning"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Strong expertise in distributed systems architecture, microservices, and event-driven architectures\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "Strong expertise in distributed systems architecture, microservices, and event-driven architectures",
"provenance": "verb-backed",
"reason_code": "out_of_role_scope",
"skill": "Microservices"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"vector databases (e.g., Milvus)\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "vector databases (e.g., Milvus)",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Milvus"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"MCP servers, custom skills, or agent frameworks\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "MCP servers, custom skills, or agent frameworks",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Model Context Protocol"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Advanced knowledge of databases (SQL/NoSQL)\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
"jd_quote": "Advanced knowledge of databases (SQL/NoSQL)",
"origin": "ai_predicted",
"provenance": "listed",
"reason_code": "not_in_catalog",
"skill": "NoSQL"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Python experience is a strong plus\" - the JD marks it as \"a plus\" rather than required.",
"jd_quote": "Python experience is a strong plus",
"provenance": "listed",
"reason_code": "jd_preferred",
"skill": "Python"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"\u2022 Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "\u2022 Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Recommender Systems"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Expert-level proficiency in Java or Scala required\" - the JD marks it as \"a plus\" rather than required.",
"jd_quote": "Expert-level proficiency in Java or Scala required",
"provenance": "verb-backed",
"reason_code": "jd_preferred",
"skill": "Scala"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"\u2022 Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
"jd_quote": "\u2022 Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale",
"provenance": "listed",
"reason_code": "out_of_role_scope",
"skill": "Vector Databases"
}
],
"secondary_skills": [
"Apache Solr",
"Machine Learning",
"Microservices",
"Milvus",
"Model Context Protocol",
"NoSQL",
"Python",
"Recommender Systems",
"Scala",
"Vector Databases"
],
"skill_clusters": [
{
"anchor": "AWS",
"kind": "interchangeable",
"layer": "L1",
"layers": {
"AWS": "L1",
"Google Cloud Platform": "L1",
"Microsoft Azure": "L1"
},
"members": [
"AWS",
"Google Cloud Platform",
"Microsoft Azure"
]
}
],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Java",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Big Data",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"name": "Apache Kafka",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"name": "AWS",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"name": "Google Cloud Platform",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"name": "Microsoft Azure",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Redpanda",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Claude Code",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Cursor",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review."
}
]
},
{
"label": "Environment",
"layer": "L2",
"skills": [
{
"dimension": {
"display_name": "Containers \u0026 Orchestration",
"slug": "containers-orchestration"
},
"name": "Docker",
"origin": "catalog",
"rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes."
},
{
"dimension": {
"display_name": "Containers \u0026 Orchestration",
"slug": "containers-orchestration"
},
"name": "Kubernetes",
"origin": "catalog",
"rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes."
},
{
"dimension": {
"display_name": "Observability \u0026 Monitoring",
"slug": "observability-monitoring"
},
"name": "Elasticsearch",
"origin": "catalog",
"rationale": "Charter owns_5 assigns run-health monitoring of owned pipelines: run metrics, log triage, and alerting on failures and SLA misses."
},
{
"dimension": {
"display_name": "Message Brokers",
"slug": "message-brokers"
},
"name": "Event-Driven Architecture",
"origin": "catalog",
"rationale": "Native by registry design: the Java catalog tiered this dimension borrowed(Data Engineer), anticipating this role as its owner."
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.