Pipeline run
9f3b184d-c078-4a0a-8764-cb146bdbb4f6
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
The Alexa Echo Device Team is seeking a Senior Data Engineer to join their Business Intelligence team, focusing on developing scalable analytical solutions within a complex data warehouse environment. The role requires expertise in Python, Scala, or Java for automating ETL processes and working with technologies such as Redshift, EMR, and Hadoop/Hive/Pig. Distinctive requirements include the ability to design complex data models and write highly optimized SQL queries over large datasets. The position involves collaboration with various teams to drive data-driven strategies for Echo and Alexa products.
Job description
The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team. We provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices. As a Senior Data Engineer, you will work in one of the world's largest and most complex data warehouse environments. You will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies. You will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up. You will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems. You will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more. You will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization. You should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms. You should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions. You should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Slowly Changing Dimensions — Data Modeling & Warehouse Design (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"company_name": "Amazon",
"ctc": null,
"domain": {
"primary": {
"aliases": [
"ITES",
"BPO"
],
"domain": "IT Services \u0026 Consulting"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"role": "Senior Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Business Intelligence Engineer",
"reasoning": "common industry alt for the same archetype",
"relation": "synonym"
},
{
"name": "ETL Engineer",
"reasoning": "role involves significant ETL responsibilities",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 0,
"heading": "Role Overview",
"heading_was_present": false,
"source_marker": {
"first_5_words": "The Alexa Echo Device Team",
"last_5_words": "Echo and Alexa devices."
},
"text": "The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team. We provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices.",
"word_count": 42
},
{
"bullet_count": 0,
"heading": "Responsibilities",
"heading_was_present": true,
"source_marker": {
"first_5_words": "As a Senior Data Engineer,",
"last_5_words": "data from production systems into"
},
"text": "As a Senior Data Engineer, you will work in one of the world\u0027s largest and most complex data warehouse environments. You will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies. You will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up. You will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems. You will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more. You will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization. You should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms. You should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions. You should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.",
"word_count": 335
}
],
"urls": []
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
},
{
"dimension": "Data Modeling \u0026 Warehouse Design",
"reason": "In the role canon\u0027s Data Modeling \u0026 Warehouse Design - adding it narrows the candidate pool.",
"skill": "Slowly Changing Dimensions",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor",
"skill_id": "eb00754a-402c-5de6-885a-a8af69e88c0a",
"skill_name": "Amazon Redshift"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "5006b0d5-d130-54c3-8635-89e045c9b7c8",
"skill_name": "Amazon EMR"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention.",
"skill_id": "6441305b-4b6f-5488-9c37-ad26d693137e",
"skill_name": "Hadoop HDFS"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention.",
"skill_id": "c75ad3dc-8eb8-5047-a1c3-a264bd1a3cbf",
"skill_name": "Apache Hive"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
"skill_name": "Python"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "8f15ce2f-7770-5f15-aeb7-7861ae881ef1",
"skill_name": "Scala"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "3f686ed1-fac0-5707-97b5-c06d4b5e8e4c",
"skill_name": "Java"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "16a8d754-3339-58e3-978c-7e033bb9b94d",
"skill_name": "ETL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review.",
"skill_id": "209fc44b-f035-5fb6-b912-edde5ceb3f55",
"skill_name": "Apache Pig"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": "Amazon",
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"industryDomain": "IT Services \u0026 Consulting",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null
},
"location": [],
"noticePeriod": null
},
"locations": [],
"noticePeriod": null,
"openToRelocate": false,
"role": "Senior Data Engineer",
"roleSynonyms": [
"Data Engineer",
"Business Intelligence Engineer",
"ETL Engineer"
]
},
"jd_summary": "The Alexa Echo Device Team is seeking a Senior Data Engineer to join their Business Intelligence team, focusing on developing scalable analytical solutions within a complex data warehouse environment. The role requires expertise in Python, Scala, or Java for automating ETL processes and working with technologies such as Redshift, EMR, and Hadoop/Hive/Pig. Distinctive requirements include the ability to design complex data models and write highly optimized SQL queries over large datasets. The position involves collaboration with various teams to drive data-driven strategies for Echo and Alexa products.",
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"company_name": "Amazon",
"ctc": null,
"domain": {
"primary": {
"aliases": [
"ITES",
"BPO"
],
"domain": "IT Services \u0026 Consulting"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"role": "Senior Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Business Intelligence Engineer",
"reasoning": "common industry alt for the same archetype",
"relation": "synonym"
},
{
"name": "ETL Engineer",
"reasoning": "role involves significant ETL responsibilities",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 0,
"heading": "Role Overview",
"heading_was_present": false,
"source_marker": {
"first_5_words": "The Alexa Echo Device Team",
"last_5_words": "Echo and Alexa devices."
},
"text": "The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team. We provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices.",
"word_count": 42
},
{
"bullet_count": 0,
"heading": "Responsibilities",
"heading_was_present": true,
"source_marker": {
"first_5_words": "As a Senior Data Engineer,",
"last_5_words": "data from production systems into"
},
"text": "As a Senior Data Engineer, you will work in one of the world\u0027s largest and most complex data warehouse environments. You will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies. You will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up. You will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems. You will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more. You will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization. You should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms. You should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions. You should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.",
"word_count": 335
}
],
"urls": []
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"role_coverage": [
{
"is_jd_role": true,
"name": "Data Engineer",
"pct": 0.41,
"role": "data-engineer",
"share": 77.4
},
{
"is_jd_role": false,
"name": "ETL Developer",
"pct": 0.3,
"role": "etl-developer",
"share": 22.6
}
],
"run_id": "jdv4-d611721e7d51",
"secondary_meta": [],
"secondary_skills": [],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Amazon EMR",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Python",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Scala",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Java",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "ETL",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"name": "Amazon Redshift",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"name": "Hadoop HDFS",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention."
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"name": "Apache Hive",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Apache Pig",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review."
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.