Pipeline run
ec62908f-b8fc-4327-8045-26d610e951e8
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
The Alexa Echo Device Team is seeking a Senior Data Engineer to join their Business Intelligence team, focusing on developing analytical solutions in a complex data warehouse environment. The role involves using technologies such as Python, Scala, Java, Redshift, EMR, and Hadoop to automate ETL processes and manage large datasets. Key requirements include expertise in designing scalable data models and writing optimized SQL queries, as well as strong interpersonal skills for collaborating with various teams to understand data needs. The position supports the development of data-driven strategies for Echo and Alexa devices.
Job description
The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team. We provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices. As a Senior Data Engineer, you will work in one of the world's largest and most complex data warehouse environments. You will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies. You will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up. You will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems. You will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more. You will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization. You should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms. You should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions. You should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Slowly Changing Dimensions — Data Modeling & Warehouse Design (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"company_name": "Amazon",
"ctc": null,
"domain": {
"primary": {
"aliases": [
"SaaS",
"Product Companies"
],
"domain": "Software \u0026 SaaS Products"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"role": "Senior Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Business Intelligence Engineer",
"reasoning": "common industry alt for the same archetype",
"relation": "synonym"
},
{
"name": "ETL Engineer",
"reasoning": "role involves significant ETL responsibilities",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 0,
"heading": "Role Overview",
"heading_was_present": false,
"source_marker": {
"first_5_words": "The Alexa Echo Device Team",
"last_5_words": "power the next generation of"
},
"text": "The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team.\nWe provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices.",
"word_count": 40
},
{
"bullet_count": 0,
"heading": "Responsibilities",
"heading_was_present": true,
"source_marker": {
"first_5_words": "As a Senior Data Engineer",
"last_5_words": "production systems into the BI"
},
"text": "As a Senior Data Engineer, you will work in one of the world\u0027s largest and most complex data warehouse environments.\nYou will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies.\nYou will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up.\nYou will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems.\nYou will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more.\nYou will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization.\nYou should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms.\nYou should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions.\nYou should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.",
"word_count": 366
}
],
"urls": []
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
},
{
"dimension": "Data Modeling \u0026 Warehouse Design",
"reason": "In the role canon\u0027s Data Modeling \u0026 Warehouse Design - adding it narrows the candidate pool.",
"skill": "Slowly Changing Dimensions",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
"skill_name": "Python"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "8f15ce2f-7770-5f15-aeb7-7861ae881ef1",
"skill_name": "Scala"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "3f686ed1-fac0-5707-97b5-c06d4b5e8e4c",
"skill_name": "Java"
},
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor",
"skill_id": "eb00754a-402c-5de6-885a-a8af69e88c0a",
"skill_name": "Amazon Redshift"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "5006b0d5-d130-54c3-8635-89e045c9b7c8",
"skill_name": "Amazon EMR"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention.",
"skill_id": "6441305b-4b6f-5488-9c37-ad26d693137e",
"skill_name": "Hadoop HDFS"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention.",
"skill_id": "c75ad3dc-8eb8-5047-a1c3-a264bd1a3cbf",
"skill_name": "Apache Hive"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "16a8d754-3339-58e3-978c-7e033bb9b94d",
"skill_name": "ETL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review.",
"skill_id": "209fc44b-f035-5fb6-b912-edde5ceb3f55",
"skill_name": "Apache Pig"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": "Amazon",
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"industryDomain": "Software \u0026 SaaS Products",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null
},
"location": [],
"noticePeriod": null
},
"locations": [],
"noticePeriod": null,
"openToRelocate": false,
"role": "Senior Data Engineer",
"roleSynonyms": [
"Data Engineer",
"Business Intelligence Engineer",
"ETL Engineer"
]
},
"jd_summary": "The Alexa Echo Device Team is seeking a Senior Data Engineer to join their Business Intelligence team, focusing on developing analytical solutions in a complex data warehouse environment. The role involves using technologies such as Python, Scala, Java, Redshift, EMR, and Hadoop to automate ETL processes and manage large datasets. Key requirements include expertise in designing scalable data models and writing optimized SQL queries, as well as strong interpersonal skills for collaborating with various teams to understand data needs. The position supports the development of data-driven strategies for Echo and Alexa devices.",
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"company_name": "Amazon",
"ctc": null,
"domain": {
"primary": {
"aliases": [
"SaaS",
"Product Companies"
],
"domain": "Software \u0026 SaaS Products"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"role": "Senior Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
},
{
"name": "Business Intelligence Engineer",
"reasoning": "common industry alt for the same archetype",
"relation": "synonym"
},
{
"name": "ETL Engineer",
"reasoning": "role involves significant ETL responsibilities",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 0,
"heading": "Role Overview",
"heading_was_present": false,
"source_marker": {
"first_5_words": "The Alexa Echo Device Team",
"last_5_words": "power the next generation of"
},
"text": "The Alexa Echo Device Team is looking for a talented, highly motivated Data Engineer to join our Business Intelligence team.\nWe provide actionable business insights that inform future products and services that will power the next generation of Echo and Alexa devices.",
"word_count": 40
},
{
"bullet_count": 0,
"heading": "Responsibilities",
"heading_was_present": true,
"source_marker": {
"first_5_words": "As a Senior Data Engineer",
"last_5_words": "production systems into the BI"
},
"text": "As a Senior Data Engineer, you will work in one of the world\u0027s largest and most complex data warehouse environments.\nYou will work closely with Product Management, Software Development, Data Science, and other Data Engineering teams to develop scalable and cutting-edge analytical solutions, process and store terabytes of low latency structured and unstructured data, and enable the Echo Device team to build successful, data driven strategies.\nYou will be responsible for designing and implementing an analytical environment using third-party and in-house tools and using Python, Scala, or Java to automate the ETL, analytics, and data quality platform from the ground up.\nYou will design and implement complex data models, model metadata, build reports and dashboards, and own data presentation and dashboarding tools for the end users of our data products and systems.\nYou will work with leading edge technologies like Redshift, EMR, Hadoop/Hive/Pig, and more.\nYou will write scalable, highly tuned SQL queries running over billions of rows of data and will develop learning and training programs to drive adoption of data driven decision making across the Echo and Alexa organization.\nYou should have deep expertise in the design, creation, management, and business use of large datasets, across a variety of data platforms.\nYou should have excellent business and interpersonal skills to be able to work with business owners to understand data requirements, and to implement efficient and scalable ETL solutions.\nYou should be an authority at crafting, implementing, and operating stable, scalable, low cost solutions to replicate data from production systems into the BI data store.",
"word_count": 366
}
],
"urls": []
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"role_coverage": [
{
"is_jd_role": true,
"name": "Data Engineer",
"pct": 0.41,
"role": "data-engineer",
"share": 77.4
},
{
"is_jd_role": false,
"name": "ETL Developer",
"pct": 0.3,
"role": "etl-developer",
"share": 22.6
}
],
"run_id": "jdv4-ddcc49e68aa7",
"secondary_meta": [],
"secondary_skills": [],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Python",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Scala",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Java",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Amazon EMR",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "ETL",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"name": "Amazon Redshift",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor"
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"name": "Hadoop HDFS",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention."
},
{
"dimension": {
"display_name": "Data Lake \u0026 Storage Formats",
"slug": "data-lake-storage-formats"
},
"name": "Apache Hive",
"origin": "catalog",
"rationale": "Charter owns_3 grants the lake storage layer outright: layout, formats, partitioning, retention."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Apache Pig",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Data Processing \u0026 Pipeline Frameworks pending admin review."
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.