← Back to history

Pipeline run

6b6a2855-db93-4d90-84f0-1954f3d8fdc7

Pipeline LLM cost (USD)
API 1: $0.0045 API 2: $0.0000 API 3: $0.0000 Total: $0.0045

Client output enrichment

v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA description
Nature of work
—
no_db_connection
Tech stack maturity
Mainstream Modern
AI index (0 = no AI use, 5 = totally AI-dependent · v2.1)
1.70 / 5
· Title match
✓ Has AI skill
✓ AI skill (primary)
· AI skill (secondary)
· On AI team
· Builds AI products
vocab breakdown (legacy)
Assistants (×1): Claude, Cursor
Frameworks (×2): Milvus
Models / concepts (×3): agentic, MCP, AI, Machine Learning
Evidence — skills matched in JD (14)
Java Apache Kafka SQL AWS Google Cloud Platform Microsoft Azure Big Data Redpanda Claude Code Cursor Docker Kubernetes Elasticsearch Event-Driven Architecture
Skill cluster (0 dimension groups, role-scoped)
No dimension groups computed for this JD.
Show KRA description ↓
• Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale • Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence • Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations • Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions • Remain significantly hands-on with critical features and architectural components • Drive technical planning, prioritization, and execution aligned with business objectives • Act as a key technical partner to the Engineering Manager in driving team success and technical decisions • Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices • Lead efforts in monitoring, observability, performance optimization, and production reliability at scale
Status: jd_deconstruction_v4_done Created: 2026-09-08T21:22:58.227467Z Updated: 2026-09-14T09:03:10.243837Z Run duration: 44317 ms
Flow v4 catalog deconstruction skill_library_v4

1 LLM mention extraction + deterministic catalog sweep

2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)

3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim

Role Chosen role & resolution

Data Engineer

in_db name family · data-engineer

slug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)

Roku is looking for a Senior Engineer to join their backend and data team, focusing on designing and optimizing distributed data pipelines and real-time processing systems. The role requires expertise in Java, distributed systems, and big data technologies, with a strong emphasis on building scalable, production-ready solutions. Candidates should have at least 8 years of software engineering experience, including a track record of leading complex projects and familiarity with streaming technologies like Kafka. Roku is committed to innovation in the TV streaming industry and values engineers who can leverage AI and automation in their work.

Job description

What does the team work on?

Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television – including TV advertising – around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers' experiences.

What is the role?

We are seeking a highly experienced and skilled Senior Engineer to join our backend and data team. This role is crucial for designing, building, and optimizing distributed data pipelines, real-time data processing systems, and backend solutions that effectively handle large-scale data. The ideal candidate will have deep expertise in Java, distributed systems, and big data technologies, and a passion for solving complex problems and delivering robust solutions. We’re always in “build mode” because we’re a company of data-focused builders. Every day, you’ll look at what exists and find ways to make it better and help drive innovation.

How will I use AI at Roku?

At Roku, we don’t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We’re looking for curious, adaptable builders who can show how they’ve used AI or automation to move faster, raise the bar, and scale their impact.
We value your AI skills if have built fluency across the agentic engineering toolchain — coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you.

What are the responsibilities of the role?
• Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale
• Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence
• Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations
• Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions
• Remain significantly hands-on with critical features and architectural components
• Drive technical planning, prioritization, and execution aligned with business objectives
• Act as a key technical partner to the Engineering Manager in driving team success and technical decisions
• Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices
• Lead efforts in monitoring, observability, performance optimization, and production reliability at scale

What experience would help someone be successful in this role at Roku?
• 8+ years of software engineering experience with significant time in senior engineering roles
• Proven expertise in building scalable, distributed, and streaming solutions in production environments
• Deep experience with content management systems, media processing, or publishing platforms
• Expert-level proficiency in Java or Scala required; Python experience is a strong plus
• Strong expertise in distributed systems architecture, microservices, and event-driven architectures
• Deep understanding of streaming technologies (Kafka, Redpanda, or similar
• Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale
• Track record of leading complex technical projects from conception to production in high-scale environments
• Excellent communication skills and ability to influence technical decisions across teams and organizations
• Experience with cloud platforms (AWS, GCP, or Azure) at enterprise scale
• Extensive experience with containerization and orchestration (Docker, Kubernetes)
• Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems

#LI-AB3

What's Roku's approach to hybrid working?

Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days' a week attendance.

What are some of the benefits?

Roku is committed to offering a diverse range of benefits as part of our compensation package to support our employees and their families. Our comprehensive benefits include global access to mental health and financial wellness support and resources. Local benefits include statutory and voluntary benefits which may include healthcare (medical, dental, and vision), life, accident, disability, commuter, and retirement options (401(k)/pension). Employees are supported in taking time off, in accordance with local leave policies and other personal needs to support their evolving work and life needs. It's important to note that not every benefit is available in all locations or for every role. For details specific to your location, please consult with your recruiter.

Accommodations

Roku welcomes applicants of all backgrounds and provides reasonable accommodations and adjustments in accordance with applicable law. If you require reasonable accommodation at any point in the hiring process, please direct your inquiries to EmployeeRelations@Roku.com.

What should I know about Roku's culture?

Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company's success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We're independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you'll be part of a company that's changing how the world watches TV. 

We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn't real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002. 

To learn more about Roku, our global footprint, and how we've grown, visit https://www.weareroku.com/factsheet.

By providing your information, you acknowledge that you want Roku to contact you about job roles, that you have read Roku's Applicant Privacy Notice, and understand that Roku will use your information as described in that notice. If you do not wish to receive any communications from Roku regarding this role or similar roles in the future, you may unsubscribe at any time by emailing WorkforcePrivacy@Roku.com.

Skills from this JD

L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.

L0 · Anchor 3 skill(s)
Java Programming Languages the craft is exercised in Python and SQL; every owned artifact — transform, DAG, quality check — is authored in them.
SQL Programming Languages the craft is exercised in Python and SQL; every owned artifact — transform, DAG, quality check — is authored in them.
Big Data Data Processing & Pipeline Frameworks transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the
L1 · Primary 7 skill(s)
Apache Kafka Data Ingestion & Integration Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;
AWS Cloud Platforms The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev
Google Cloud Platform Cloud Platforms The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev
Microsoft Azure Cloud Platforms The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev
Redpanda AI GENERATED Data Processing & Pipeline Frameworks TaBuddy AI prediction — not in the skill library yet; placed under Data Processing & Pipeline Frameworks pending admin review.
Claude Code AI GENERATED Programming Languages TaBuddy AI prediction — not in the skill library yet; placed under Programming Languages pending admin review.
Cursor AI GENERATED Programming Languages TaBuddy AI prediction — not in the skill library yet; placed under Programming Languages pending admin review.
L2 · Environment 4 skill(s)
Docker Containers & Orchestration Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.
Kubernetes Containers & Orchestration Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.
Elasticsearch Observability & Monitoring Charter owns_5 assigns run-health monitoring of owned pipelines: run metrics, log triage, and alerting on failures and SLA misses.
Event-Driven Architecture Message Brokers Native by registry design: the Java catalog tiered this dimension borrowed(Data Engineer), anticipating this role as its owner.
Secondary nice-to-have · not a drop-order layer
Apache Solr out_of_role_scope

This skill is in Secondary as the JD says "Experience with search technologies (Elasticsearch, Solr, etc)" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Machine Learning out_of_role_scope

This skill is in Secondary as the JD says "While we are well-positioned to help shape the future of television – including TV advertising – around the world, continued success relies on its investment in our machine learning capabilities." - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Microservices out_of_role_scope

This skill is in Secondary as the JD says "Strong expertise in distributed systems architecture, microservices" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Milvus out_of_role_scope

This skill is in Secondary as the JD says "vector databases (e.g., Milvus)" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Model Context Protocol out_of_role_scope

This skill is in Secondary as the JD says "MCP servers, custom skills, or agent frameworks" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

NoSQL AI GENERATED not_in_catalog

This skill is in Secondary as the JD says "Advanced knowledge of databases (SQL/NoSQL)" - it isn't in this role's skill catalog yet, so it's tracked for review.

Python jd_preferred

This skill is in Secondary as the JD says "Python experience is a strong plus" - the JD marks it as "a plus" rather than required.

Recommender Systems out_of_role_scope

This skill is in Secondary as the JD says "• Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Scala jd_preferred

This skill is in Secondary as the JD says "Expert-level proficiency in Java or Scala required" - the JD marks it as "a plus" rather than required.

Vector Databases out_of_role_scope

This skill is in Secondary as the JD says "• Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale" - it's a recognized skill, but it sits outside the Data Engineer role's core dimensions.

Consider adding canon exemplars the JD missed

Apache Beam — Data Processing & Pipeline Frameworks (L0)

Prefect — Workflow Orchestration (L0)

Google Cloud Storage — Cloud Platforms (L1)

Airbyte — Data Ingestion & Integration (L1)

Apache Parquet — Data Lake & Storage Formats (L1)

Slowly Changing Dimensions — Data Modeling & Warehouse Design (L1)

Library artifacts (this run)

No artifact rows for this run.
nano JD Parser — gpt-4.1-nano click to toggle
RoleSenior Engineer
CompanyRoku
Experience8+ years of software engineering experience
DomainSoftware & SaaS Products
Location — (hybrid)
JD type pass
Show raw JSON
{
  "JD_type": "pass",
  "about_company": {
    "source_marker": {
      "first_5_words": "Roku pioneered streaming to the",
      "last_5_words": "to shape customers\u0027 experiences."
    },
    "text": "Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers\u0027 experiences.",
    "word_count": 84
  },
  "ai_kras": [
    "We value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks.",
    "You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you."
  ],
  "certifications": [],
  "company_name": "Roku",
  "ctc": null,
  "domain": {
    "primary": {
      "aliases": [
        "SaaS",
        "Product Companies"
      ],
      "domain": "Software \u0026 SaaS Products"
    },
    "secondary": null
  },
  "education": [],
  "experience": {
    "max": null,
    "min": 8,
    "raw": "8+ years of software engineering experience"
  },
  "job_locations": [
    {
      "aliases": [],
      "city": null,
      "country": null,
      "state": null,
      "work_mode": "hybrid"
    }
  ],
  "role": "Senior Engineer",
  "role_aliases": [
    {
      "name": "Software Engineer",
      "reasoning": "common industry alt for the same archetype",
      "relation": "synonym"
    },
    {
      "name": "Backend Engineer",
      "reasoning": "role focuses on backend solutions and data pipelines",
      "relation": "synonym"
    },
    {
      "name": "Data Engineer",
      "reasoning": "role involves building and optimizing data pipelines",
      "relation": "synonym"
    }
  ],
  "role_archetype": "Engineering",
  "roles_and_responsibilities": [
    {
      "bullet_count": 9,
      "heading": "Responsibilities of the role",
      "heading_was_present": true,
      "source_marker": {
        "first_5_words": "\u2022 Define architecture and technical",
        "last_5_words": "and production reliability at scale"
      },
      "text": "\u2022 Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale\n\u2022 Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence\n\u2022 Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations\n\u2022 Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions\n\u2022 Remain significantly hands-on with critical features and architectural components\n\u2022 Drive technical planning, prioritization, and execution aligned with business objectives\n\u2022 Act as a key technical partner to the Engineering Manager in driving team success and technical decisions\n\u2022 Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices\n\u2022 Lead efforts in monitoring, observability, performance optimization, and production reliability at scale",
      "word_count": 174
    }
  ],
  "urls": [
    {
      "type": "website",
      "url": "https://www.weareroku.com/factsheet"
    }
  ]
}
API 1 — extract-from-jd click to toggle
{
  "catalog": "v4",
  "consider_adding": [
    {
      "dimension": "Data Processing \u0026 Pipeline Frameworks",
      "reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
      "skill": "Apache Beam",
      "tier": "L0"
    },
    {
      "dimension": "Workflow Orchestration",
      "reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
      "skill": "Prefect",
      "tier": "L0"
    },
    {
      "dimension": "Cloud Platforms",
      "reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
      "skill": "Google Cloud Storage",
      "tier": "L1"
    },
    {
      "dimension": "Data Ingestion \u0026 Integration",
      "reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
      "skill": "Airbyte",
      "tier": "L1"
    },
    {
      "dimension": "Data Lake \u0026 Storage Formats",
      "reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
      "skill": "Apache Parquet",
      "tier": "L1"
    },
    {
      "dimension": "Data Modeling \u0026 Warehouse Design",
      "reason": "In the role canon\u0027s Data Modeling \u0026 Warehouse Design - adding it narrows the candidate pool.",
      "skill": "Slowly Changing Dimensions",
      "tier": "L1"
    }
  ],
  "final_skills": [
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
      "skill_id": "3f686ed1-fac0-5707-97b5-c06d4b5e8e4c",
      "skill_name": "Java"
    },
    {
      "dimension": {
        "display_name": "Data Ingestion \u0026 Integration",
        "slug": "data-ingestion-integration"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;",
      "skill_id": "a2275eea-20e3-59dd-b201-af76b7ae706f",
      "skill_name": "Apache Kafka"
    },
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
      "skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
      "skill_name": "SQL"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
      "skill_id": "78c5a4b9-2ab5-5279-a688-46f4c3767c3f",
      "skill_name": "AWS"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
      "skill_id": "de9d33d1-f608-5b2a-97b7-7398054f33cc",
      "skill_name": "Google Cloud Platform"
    },
    {
      "dimension": {
        "display_name": "Cloud Platforms",
        "slug": "cloud-platforms"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
      "skill_id": "060b1d39-d4d8-581d-9d96-65cd45d7cea4",
      "skill_name": "Microsoft Azure"
    },
    {
      "dimension": {
        "display_name": "Containers \u0026 Orchestration",
        "slug": "containers-orchestration"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.",
      "skill_id": "37ac0455-c683-5062-9393-96c9284f5ed1",
      "skill_name": "Docker"
    },
    {
      "dimension": {
        "display_name": "Containers \u0026 Orchestration",
        "slug": "containers-orchestration"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes.",
      "skill_id": "57e0c6c2-06e3-5d5a-a0cd-249da956a7e3",
      "skill_name": "Kubernetes"
    },
    {
      "dimension": {
        "display_name": "Observability \u0026 Monitoring",
        "slug": "observability-monitoring"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Charter owns_5 assigns run-health monitoring of owned pipelines: run metrics, log triage, and alerting on failures and SLA misses.",
      "skill_id": "48406a13-f524-5e91-af33-a3ef313fe54d",
      "skill_name": "Elasticsearch"
    },
    {
      "dimension": {
        "display_name": "Message Brokers",
        "slug": "message-brokers"
      },
      "is_primary": false,
      "layer": "L2",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "Native by registry design: the Java catalog tiered this dimension borrowed(Data Engineer), anticipating this role as its owner.",
      "skill_id": "456197d8-7e5d-5f86-9069-411b8eafe2c2",
      "skill_name": "Event-Driven Architecture"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "catalog_v4",
      "origin": "catalog",
      "rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
      "skill_id": "dbe274b4-5b86-599d-bac5-1fc85691053a",
      "skill_name": "Big Data"
    },
    {
      "dimension": {
        "display_name": "Data Processing \u0026 Pipeline Frameworks",
        "slug": "data-processing-pipeline-frameworks"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "ai_predicted",
      "origin": "ai_predicted",
      "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review.",
      "skill_id": "08775033-5e69-55de-b80d-518b6143050b",
      "skill_name": "Redpanda"
    },
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "ai_predicted",
      "origin": "ai_predicted",
      "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review.",
      "skill_id": "f683d14a-c143-528d-92d6-f4bca348cb58",
      "skill_name": "Claude Code"
    },
    {
      "dimension": {
        "display_name": "Programming Languages",
        "slug": "programming-languages"
      },
      "is_primary": true,
      "layer": "L1",
      "layer_source": "ai_predicted",
      "origin": "ai_predicted",
      "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review.",
      "skill_id": "82273e02-f063-501f-ac18-777e95151380",
      "skill_name": "Cursor"
    }
  ],
  "history_run_id": null,
  "jd_parameters": {
    "certifications": [],
    "clientDetails": null,
    "company": "Roku",
    "companySize": null,
    "ctc": {
      "currency": null,
      "max": null,
      "min": null,
      "period": null,
      "raw": null
    },
    "educationRequirements": [],
    "experience": {
      "max": null,
      "min": 8,
      "raw": "8+ years of software engineering experience"
    },
    "industryDomain": "Software \u0026 SaaS Products",
    "knockouts": {
      "certifications": [],
      "ctc": {
        "currency": null,
        "max": null,
        "min": null
      },
      "educationRequirements": [],
      "experience": {
        "max": null,
        "min": 8
      },
      "location": [
        "Hybrid"
      ],
      "noticePeriod": null
    },
    "locations": [
      "Hybrid"
    ],
    "noticePeriod": null,
    "openToRelocate": false,
    "role": "Senior Engineer",
    "roleSynonyms": [
      "Software Engineer",
      "Backend Engineer",
      "Data Engineer"
    ]
  },
  "jd_summary": "Roku is looking for a Senior Engineer to join their backend and data team, focusing on designing and optimizing distributed data pipelines and real-time processing systems. The role requires expertise in Java, distributed systems, and big data technologies, with a strong emphasis on building scalable, production-ready solutions. Candidates should have at least 8 years of software engineering experience, including a track record of leading complex projects and familiarity with streaming technologies like Kafka. Roku is committed to innovation in the TV streaming industry and values engineers who can leverage AI and automation in their work.",
  "layer_conflicts": [],
  "nano_parsed": {
    "JD_type": "pass",
    "about_company": {
      "source_marker": {
        "first_5_words": "Roku pioneered streaming to the",
        "last_5_words": "to shape customers\u0027 experiences."
      },
      "text": "Roku pioneered streaming to the TV and continues to innovate and lead the industry. While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities. Roku offers millions of options to our users: movies, episodes, news, sports, and channels from all around the world. The Roku Content Platform is key to onboarding content into the Roku ecosystem, delighting our customers. We are building a content knowledge platform that provides insights to downstream systems like Search, Recommendations, Ads, and Voice to shape customers\u0027 experiences.",
      "word_count": 84
    },
    "ai_kras": [
      "We value your AI skills if have built fluency across the agentic engineering toolchain \u2014 coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks.",
      "You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you."
    ],
    "certifications": [],
    "company_name": "Roku",
    "ctc": null,
    "domain": {
      "primary": {
        "aliases": [
          "SaaS",
          "Product Companies"
        ],
        "domain": "Software \u0026 SaaS Products"
      },
      "secondary": null
    },
    "education": [],
    "experience": {
      "max": null,
      "min": 8,
      "raw": "8+ years of software engineering experience"
    },
    "job_locations": [
      {
        "aliases": [],
        "city": null,
        "country": null,
        "state": null,
        "work_mode": "hybrid"
      }
    ],
    "role": "Senior Engineer",
    "role_aliases": [
      {
        "name": "Software Engineer",
        "reasoning": "common industry alt for the same archetype",
        "relation": "synonym"
      },
      {
        "name": "Backend Engineer",
        "reasoning": "role focuses on backend solutions and data pipelines",
        "relation": "synonym"
      },
      {
        "name": "Data Engineer",
        "reasoning": "role involves building and optimizing data pipelines",
        "relation": "synonym"
      }
    ],
    "role_archetype": "Engineering",
    "roles_and_responsibilities": [
      {
        "bullet_count": 9,
        "heading": "Responsibilities of the role",
        "heading_was_present": true,
        "source_marker": {
          "first_5_words": "\u2022 Define architecture and technical",
          "last_5_words": "and production reliability at scale"
        },
        "text": "\u2022 Define architecture and technical strategy, ensuring scalability, reliability, and performance at scale\n\u2022 Provide technical guidance, conduct code reviews, and mentor engineers to elevate team capabilities and foster engineering excellence\n\u2022 Design and implement robust distributed systems, streaming solutions, content management systems, APIs, and data pipelines that handle high-volume content operations\n\u2022 Partner with Product, Operations, Business, and other engineering teams to deliver integrated solutions\n\u2022 Remain significantly hands-on with critical features and architectural components\n\u2022 Drive technical planning, prioritization, and execution aligned with business objectives\n\u2022 Act as a key technical partner to the Engineering Manager in driving team success and technical decisions\n\u2022 Champion a culture of innovation, technical excellence, and continuous improvement; establish engineering best practices\n\u2022 Lead efforts in monitoring, observability, performance optimization, and production reliability at scale",
        "word_count": 174
      }
    ],
    "urls": [
      {
        "type": "website",
        "url": "https://www.weareroku.com/factsheet"
      }
    ]
  },
  "pipeline": "v4",
  "rejected": false,
  "rejection_code": null,
  "rejection_reason": null,
  "role": {
    "canonical_name": "Data Engineer",
    "family": "data-engineer",
    "match_method": "name",
    "resolution": "in_db",
    "role_id": "1839a54d-e909-543e-b36e-745551d4e000",
    "similarity": null,
    "slug": "data-engineer"
  },
  "role_coverage": [
    {
      "is_jd_role": true,
      "name": "Data Engineer",
      "pct": 0.37,
      "role": "data-engineer",
      "share": 100.0
    }
  ],
  "run_id": "jdv4-4d0628c6cbe0",
  "secondary_meta": [
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Experience with search technologies (Elasticsearch, Solr, etc)\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "Experience with search technologies (Elasticsearch, Solr, etc)",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Apache Solr"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities.\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "While we are well-positioned to help shape the future of television \u2013 including TV advertising \u2013 around the world, continued success relies on its investment in our machine learning capabilities.",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Machine Learning"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Strong expertise in distributed systems architecture, microservices\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "Strong expertise in distributed systems architecture, microservices",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Microservices"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"vector databases (e.g., Milvus)\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "vector databases (e.g., Milvus)",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Milvus"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"MCP servers, custom skills, or agent frameworks\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "MCP servers, custom skills, or agent frameworks",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Model Context Protocol"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Advanced knowledge of databases (SQL/NoSQL)\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
      "jd_quote": "Advanced knowledge of databases (SQL/NoSQL)",
      "origin": "ai_predicted",
      "provenance": "listed",
      "reason_code": "not_in_catalog",
      "skill": "NoSQL"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Python experience is a strong plus\" - the JD marks it as \"a plus\" rather than required.",
      "jd_quote": "Python experience is a strong plus",
      "provenance": "listed",
      "reason_code": "jd_preferred",
      "skill": "Python"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"\u2022 Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "\u2022 Experience with search technologies (Elasticsearch, Solr, etc) or recommendation systems",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Recommender Systems"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"Expert-level proficiency in Java or Scala required\" - the JD marks it as \"a plus\" rather than required.",
      "jd_quote": "Expert-level proficiency in Java or Scala required",
      "provenance": "listed",
      "reason_code": "jd_preferred",
      "skill": "Scala"
    },
    {
      "audit_reasoning": "This skill is in Secondary as the JD says \"\u2022 Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale\" - it\u0027s a recognized skill, but it sits outside the Data Engineer role\u0027s core dimensions.",
      "jd_quote": "\u2022 Advanced knowledge of databases (SQL/NoSQL), vector databases (e.g., Milvus), caching strategies, and data modeling at scale",
      "provenance": "listed",
      "reason_code": "out_of_role_scope",
      "skill": "Vector Databases"
    }
  ],
  "secondary_skills": [
    "Apache Solr",
    "Machine Learning",
    "Microservices",
    "Milvus",
    "Model Context Protocol",
    "NoSQL",
    "Python",
    "Recommender Systems",
    "Scala",
    "Vector Databases"
  ],
  "skill_clusters": [
    {
      "anchor": "Microsoft Azure",
      "kind": "interchangeable",
      "layer": "L1",
      "layers": {
        "AWS": "L1",
        "Google Cloud Platform": "L1",
        "Microsoft Azure": "L1"
      },
      "members": [
        "AWS",
        "Google Cloud Platform",
        "Microsoft Azure"
      ]
    }
  ],
  "skill_layers": [
    {
      "label": "Anchor",
      "layer": "L0",
      "skills": [
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "Java",
          "origin": "catalog",
          "rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
        },
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "SQL",
          "origin": "catalog",
          "rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Big Data",
          "origin": "catalog",
          "rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
        }
      ]
    },
    {
      "label": "Primary",
      "layer": "L1",
      "skills": [
        {
          "dimension": {
            "display_name": "Data Ingestion \u0026 Integration",
            "slug": "data-ingestion-integration"
          },
          "name": "Apache Kafka",
          "origin": "catalog",
          "rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;"
        },
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "AWS",
          "origin": "catalog",
          "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
        },
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "Google Cloud Platform",
          "origin": "catalog",
          "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
        },
        {
          "dimension": {
            "display_name": "Cloud Platforms",
            "slug": "cloud-platforms"
          },
          "name": "Microsoft Azure",
          "origin": "catalog",
          "rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
        },
        {
          "dimension": {
            "display_name": "Data Processing \u0026 Pipeline Frameworks",
            "slug": "data-processing-pipeline-frameworks"
          },
          "name": "Redpanda",
          "origin": "ai_predicted",
          "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review."
        },
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "Claude Code",
          "origin": "ai_predicted",
          "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review."
        },
        {
          "dimension": {
            "display_name": "Programming Languages",
            "slug": "programming-languages"
          },
          "name": "Cursor",
          "origin": "ai_predicted",
          "rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Programming Languages pending admin review."
        }
      ]
    },
    {
      "label": "Environment",
      "layer": "L2",
      "skills": [
        {
          "dimension": {
            "display_name": "Containers \u0026 Orchestration",
            "slug": "containers-orchestration"
          },
          "name": "Docker",
          "origin": "catalog",
          "rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes."
        },
        {
          "dimension": {
            "display_name": "Containers \u0026 Orchestration",
            "slug": "containers-orchestration"
          },
          "name": "Kubernetes",
          "origin": "catalog",
          "rationale": "Pipeline workloads ship as containers: Airflow on Kubernetes, Spark-on-K8s, and containerised jobs are default deployment shapes."
        },
        {
          "dimension": {
            "display_name": "Observability \u0026 Monitoring",
            "slug": "observability-monitoring"
          },
          "name": "Elasticsearch",
          "origin": "catalog",
          "rationale": "Charter owns_5 assigns run-health monitoring of owned pipelines: run metrics, log triage, and alerting on failures and SLA misses."
        },
        {
          "dimension": {
            "display_name": "Message Brokers",
            "slug": "message-brokers"
          },
          "name": "Event-Driven Architecture",
          "origin": "catalog",
          "rationale": "Native by registry design: the Java catalog tiered this dimension borrowed(Data Engineer), anticipating this role as its owner."
        }
      ]
    }
  ],
  "unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}

LLM Calls

Every model call made for this run, in pipeline order. Click a card to see the model's response.

Loading…