← Back to history

Pipeline run

c9eddb74-c671-4d76-b490-b9d9b04cd64e

Pipeline LLM cost (USD)
API 1: $0.0032 API 2: $0.0002 API 3: $0.0000 Total: $0.0034

Client output enrichment

v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA description
role baseline loaded sources · ai_index: jd · nature_of_work: jd · tech_stack_maturity: jd
Nature of work · Data pipeline development
Build and run Azure Databricks/ADF ETL pipelines in PySpark and Spark SQL, maintaining ADLS Lakehouse flows with cleansing, deduplication, and production monitoring/troubleshooting. Partner with BI, Data Science, and DevOps teams on enterprise analytics delivery.
"Design and develop scalable data pipelines using PySpark & Spark SQL on Azure Databricks"
Tech stack maturity
Mainstream Legacy cache hit
A data engineer role centered on SQL typically maps to well-established, widely used data workflows and tooling rather than bleeding-edge or cloud-native specialization.
AI index (0 = no AI use, 5 = totally AI-dependent · v2.1)
0.20 / 5
· Title match
Has AI skill
· AI skill (primary)
· AI skill (secondary)
· On AI team
· Builds AI products
vocab breakdown (legacy)
Assistants (×1):
Frameworks (×2):
Models / concepts (×3): MLOps, AI, Generative AI
Evidence — skills matched in JD (7)
PySpark Spark SQL Azure Databricks Azure Data Factory ADLS SQL ETL
Skill cluster (2 dimension groups, role-scoped)
Programming Languages for Data Work
SQL
Cross-cutting / unaligned
PySpark Spark SQL Azure Databricks Azure Data Factory ADLS ETL
Show KRA description ↓
1. Design and develop scalable data pipelines using PySpark & • Spark SQL on Azure Databricks 2. Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3. Develop and maintain Lakehouse architecture using ADLS & • Databricks 4. Perform data cleansing, normalization, deduplication, and transformations 5. Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6. Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives 1. 5 years of experience in Data Engineering 2. Strong hands-on experience with Azure Databricks, Azure Data Factory, ADLS 3. Expertise in PySpark, Spark SQL, SQL, ETL Development 4. Good understanding of data architecture and cloud platforms 5. Robust communication and stakeholder management skills 6. Immediate to short notice candidates preferred

Signals

Skill full-stack-engineer
0.14
Alias data-engineer
1.00
KRA data-engineer
0.66

Post-classification

Centroidupdated · n=526
Alias collision log
New-role queue
New skills captured6
New KRA captured

Captured for admin review

PySpark primary Data Engineer pending
Spark SQL primary Data Engineer pending
Azure Databricks primary Data Engineer pending
Azure Data Factory primary Data Engineer pending
ADLS primary Data Engineer pending
ETL primary Data Engineer pending
Status: completed Created: 2026-07-26T19:33:13.866098Z Updated: 2026-07-27T10:49:21.406444Z API 3 duration: 110 ms
Flow Current 3-step pipeline

1 POST /skills/extract-from-jd

2 POST /skills/extract-details

3 POST /skills/final-role-output

Role Chosen role & resolution

Data Engineer

CASE A

slug: data-engineer · id: 2 · source: db

Exact alias hit on data-engineer (1.0) — no other alias at this confidence; skill_top full-stack-engineer 0.14 and KRA_top data-engineer 0.66 do not contradict

Resolution: in_db — role exists in library; skill↔dim and role↔dim links saved when applicable.

0
New skills
0
Skill↔dim saved
0
Role↔dim saved
1
Skipped

Job description

Exciting opportunity for an experienced Data Engineer (Databricks) to join a leading global AI &
• Analytics organization. Looking for professionals with robust expertise in Azure Databricks, PySpark, ADF, SQL, and ETL pipelines to build scalable enterprise data solutions.

Location: Bengaluru / Hyderabad / Gurugram Your Future Employer: Our client is a globally recognized AI &
• Analytics company known for delivering cutting-edge data solutions, advanced analytics, data engineering, governance, MLOps, and Generative AI innovations to leading organizations worldwide.

Responsibilities: 1. Design and develop scalable data pipelines using PySpark &
• Spark SQL on Azure Databricks 2. Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3. Develop and maintain Lakehouse architecture using ADLS &
• Databricks 4. Perform data cleansing, normalization, deduplication, and transformations 5. Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6. Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives Requirements: 1. 5 years of experience in Data Engineering 2. Strong hands-on experience with Azure Databricks, Azure Data Factory, ADLS 3. Expertise in PySpark, Spark SQL, SQL, ETL Development 4. Good understanding of data architecture and cloud platforms 5. Robust communication and stakeholder management skills 6. Immediate to short notice candidates preferred What is in it for you - 1. Opportunity to work with a leading AI &
• Analytics brand 2. Exposure to enterprise-scale cloud data projects 3. High-growth learning environment with latest technologies Reach Us :If you think this role is aligned with your career, kindly write me an email along with your updated CV on for a confidential discussion on the role.

Skills from this JD

Each row merges API 1 extraction, API 2 library match / v3 orchestration (dimensions + locked dims), and API 3 persistence tags.

PySpark Primary Library skill API 3: existing canonical (in_db) Existing skill (matched library)
Canonical: Apache Spark id=1350 · apache-spark

Aliases — catalog

  • Apache Spark (CANONICAL)
  • apache spark 3 (VERSION)
  • spark (VERSION)
  • spark 3 (VERSION)
  • spark 3.x (VERSION)
  • spark3 (VERSION)

Context tags (catalog)

Apache Kafka Cluster Manager DAGScheduler Data Lake DataFrame ETL Hadoop MLlib Machine Learning PySpark RDD Scala Spark SQL Spark Streaming SparkSession

Stored enrichment (catalog DB)

Category
Framework
Sub-category
Distributed Data Processing Framework
Vendor
Apache Software Foundation
License
apache_2
Year introduced
2010
Confidence
0.94
Version strategy
SEPARATE_ENTITY
Version tag
3.x

Maturity reasoning: Apache Spark appears in many data engineering JDs and remains a standard for distributed ETL/ELT; its GitHub and vendor ecosystem activity stay strong, with Databricks and cloud platforms still promoting it.

Skill profile (library / DB)

Skill nature
FRAMEWORK
Volatility
STABLE
Typical lifespan
EVERGREEN
Category id
5
Sub-category id
1021
Extractable
True
Also category
False

Dimensions (API 2 worklist)

  • ETL and ELT Tooling Catalog dimension db id 24

    Library dimension (catalog)

    Roles linked in library: Data Engineer

API 3 link attempts (this skill)

Dimension Skill↔dim Role↔dim Outcome
ETL and ELT Tooling
etl-and-elt-tooling
Skipped — no persistable v3 meta for new skill
skill_not_in_db_v3_proposed
Spark SQL Primary New / orchestrated API 3: new canonical path (new) New / unmatched skill (orchestrated in API 2)

Skill enrichment (orchestrator / LLM)

No Stage 7 enrichment blob on this skill (orchestrator skipped enrichment).

Derived legacy fields
Category
Data Engineering Tools
Sub-category
general
Skill nature
LANGUAGE
Volatility
MEDIUM
Typical lifespan
MULTI_YEAR
Version strategy
UNVERSIONED
Azure Databricks Primary New / orchestrated API 3: new canonical path (new) New / unmatched skill (orchestrated in API 2)

Skill enrichment (orchestrator / LLM)

No Stage 7 enrichment blob on this skill (orchestrator skipped enrichment).

Derived legacy fields
Category
Cloud Platforms
Sub-category
general
Skill nature
PLATFORM
Volatility
MEDIUM
Typical lifespan
MULTI_YEAR
Version strategy
UNVERSIONED
Azure Data Factory Primary New / orchestrated API 3: new canonical path (new) New / unmatched skill (orchestrated in API 2)

Skill enrichment (orchestrator / LLM)

No Stage 7 enrichment blob on this skill (orchestrator skipped enrichment).

Derived legacy fields
Category
Cloud Platforms
Sub-category
general
Skill nature
PLATFORM
Volatility
MEDIUM
Typical lifespan
MULTI_YEAR
Version strategy
UNVERSIONED
ADLS Primary New / orchestrated API 3: new canonical path (new) New / unmatched skill (orchestrated in API 2)

Skill enrichment (orchestrator / LLM)

No Stage 7 enrichment blob on this skill (orchestrator skipped enrichment).

Derived legacy fields
Category
Cloud Platforms
Sub-category
general
Skill nature
PLATFORM
Volatility
MEDIUM
Typical lifespan
MULTI_YEAR
Version strategy
UNVERSIONED
SQL Primary Library skill API 3: existing canonical (in_db) Existing skill (matched library)
Canonical: SQL id=101 · sql

Aliases — catalog

  • SQL (CANONICAL) primary

Context tags (catalog)

ACID CTE DDL DML ETL JOIN MySQL NoSQL OLAP ORM PostgreSQL SQL injection SQLite T-SQL data modeling data warehousing database normalization execution plan indexing joins normalization query optimization stored procedures subquery transaction isolation transaction management window functions

Stored enrichment (catalog DB)

Category
Language
Sub-category
Query Language
Vendor
ISO/IEC
License
unknown
Year introduced
1986
Confidence
0.99
Version strategy
NOT_APPLICABLE

Maturity reasoning: SQL is a hiring-pipeline staple across data, backend, and analytics roles; it appears in a very high volume of job descriptions and remains the standard query language for relational databases.

Skill profile (library / DB)

Skill nature
LANGUAGE
Volatility
STABLE
Typical lifespan
EVERGREEN
Category id
6
Sub-category id
97
Extractable
True
Also category
False

Dimensions (API 2 worklist)

  • Pega Programming Languages & DSLs Catalog dimension db id 267

    Library dimension (catalog)

    Roles linked in library: Pega Developer

  • Programming Languages Catalog dimension db id 1

    Library dimension (catalog)

    Roles linked in library: Backend Developer, Fullstack Developer, Fullstack Developer, WordPress Dev

  • Programming Languages & DSLs Catalog dimension db id 475

    Library dimension (catalog)

    Roles linked in library: Engineering Manager

  • Programming Languages for Data Work Catalog dimension db id 21

    Library dimension (catalog)

    Roles linked in library: Data Engineer

API 3 link attempts (this skill)

Dimension Skill↔dim Role↔dim Outcome
Pega Programming Languages & DSLs
pega-programming-languages-dsls
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
Programming Languages
programming-languages
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
Programming Languages & DSLs
programming-languages-dsls
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
Programming Languages for Data Work
programming-languages-for-data-work
Existing dimension (library) · Role↔dimension saved
ETL Primary New / orchestrated API 3: new canonical path (new) New / unmatched skill (orchestrated in API 2)

Skill enrichment (orchestrator / LLM)

No Stage 7 enrichment blob on this skill (orchestrator skipped enrichment).

Derived legacy fields
Category
Data Engineering Tools
Sub-category
general
Skill nature
PRACTICE
Volatility
MEDIUM
Typical lifespan
MULTI_YEAR
Version strategy
UNVERSIONED

All API 3 persistence rows

Same grid as the skill-extractor “Persistence items” table: one row per (skill × dimension) work item.

Skill Tag Dimension Skill↔dim Role↔dim Outcome Notes
PySpark new
ETL and ELT Tooling
etl-and-elt-tooling
Skipped — no persistable v3 meta for new skill skill_not_in_db_v3_proposed
SQL in_db
Pega Programming Languages & DSLs
pega-programming-languages-dsls
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
SQL in_db
Programming Languages
programming-languages
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
SQL in_db
Programming Languages & DSLs
programming-languages-dsls
Existing dimension (library) · Role↔dimension skipped (dimension not under chosen role)
SQL in_db
Programming Languages for Data Work
programming-languages-for-data-work
Existing dimension (library) · Role↔dimension saved

Library artifacts (this run)

Kind Detail DB id
canonical_skill_proposed Spark SQL | type=Data Engineering Tools subtype=general nature=LANGUAGE lifespan=MULTI_YEAR
canonical_skill_proposed Azure Databricks | type=Cloud Platforms subtype=general nature=PLATFORM lifespan=MULTI_YEAR
canonical_skill_proposed Azure Data Factory | type=Cloud Platforms subtype=general nature=PLATFORM lifespan=MULTI_YEAR
canonical_skill_proposed ADLS | type=Cloud Platforms subtype=general nature=PLATFORM lifespan=MULTI_YEAR
canonical_skill_proposed ETL | type=Data Engineering Tools subtype=general nature=PRACTICE lifespan=MULTI_YEAR
dimension_skill_link_proposed PySpark ↔ ETL and ELT Tooling
role_dimension_link_proposed Data Engineer ↔ ETL and ELT Tooling
nano JD Parser — gpt-4.1-nano click to toggle
RoleData Engineer (Databricks)
CompanyOur client
Experience5 years of experience in Data Engineering
DomainIT Services & Consulting
Location Bengaluru, India (null)
JD type pass
Show raw JSON
{
  "JD_type": "pass",
  "about_company": {
    "source_marker": {
      "first_5_words": "Our client is a globally",
      "last_5_words": "organizations worldwide."
    },
    "text": "Our client is a globally recognized AI \u0026\n\u2022 Analytics company known for delivering cutting-edge data solutions, advanced analytics, data engineering, governance, MLOps, and Generative AI innovations to leading organizations worldwide.",
    "word_count": 40
  },
  "ai_kras": [],
  "certifications": [],
  "company_name": "Our client",
  "ctc": null,
  "domain": {
    "primary": {
      "aliases": [
        "ITES",
        "BPO"
      ],
      "domain": "IT Services \u0026 Consulting"
    },
    "secondary": null
  },
  "education": [],
  "experience": {
    "max": null,
    "min": 5,
    "raw": "5 years of experience in Data Engineering"
  },
  "job_locations": [
    {
      "aliases": [
        "Bangalore"
      ],
      "city": "Bengaluru",
      "country": "India",
      "state": "Karnataka",
      "work_mode": "null"
    },
    {
      "aliases": [],
      "city": "Hyderabad",
      "country": "India",
      "state": "Telangana",
      "work_mode": "null"
    },
    {
      "aliases": [
        "Gurgaon"
      ],
      "city": "Gurugram",
      "country": "India",
      "state": "Haryana",
      "work_mode": "null"
    }
  ],
  "role": "Data Engineer (Databricks)",
  "role_aliases": [
    {
      "name": "Data Engineer",
      "reasoning": "generalized form of the picked role",
      "relation": "synonym"
    },
    {
      "name": "Data Engineer (Azure)",
      "reasoning": "specific focus on Azure technologies",
      "relation": "synonym"
    },
    {
      "name": "ETL Developer",
      "reasoning": "role involves significant ETL development",
      "relation": "synonym"
    }
  ],
  "role_archetype": "Data",
  "roles_and_responsibilities": [
    {
      "bullet_count": 6,
      "heading": "Responsibilities",
      "heading_was_present": true,
      "source_marker": {
        "first_5_words": "1. Design and develop scalable",
        "last_5_words": "and DevOps teams on enterprise"
      },
      "text": "1. Design and develop scalable data pipelines using PySpark \u0026\n\u2022 Spark SQL on Azure Databricks 2. Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3. Develop and maintain Lakehouse architecture using ADLS \u0026\n\u2022 Databricks 4. Perform data cleansing, normalization, deduplication, and transformations 5. Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6. Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives",
      "word_count": 66
    },
    {
      "bullet_count": 6,
      "heading": "Requirements",
      "heading_was_present": true,
      "source_marker": {
        "first_5_words": "1. 5 years of experience",
        "last_5_words": "short notice candidates preferred"
      },
      "text": "1. 5 years of experience in Data Engineering 2. Strong hands-on experience with Azure Databricks, Azure Data Factory, ADLS 3. Expertise in PySpark, Spark SQL, SQL, ETL Development 4. Good understanding of data architecture and cloud platforms 5. Robust communication and stakeholder management skills 6. Immediate to short notice candidates preferred",
      "word_count": 66
    }
  ],
  "urls": []
}
API 1 — extract-from-jd click to toggle
{
  "final_skills": [
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is backed by a responsibility that involves designing and developing scalable data pipelines using PySpark.",
      "skill_name": "PySpark"
    },
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is mentioned in the context of building and orchestrating ETL workflows, indicating its primary importance in the role.",
      "skill_name": "Spark SQL"
    },
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is referenced in relation to Spark SQL and ETL workflows, highlighting its primary role in the data engineering tasks.",
      "skill_name": "Azure Databricks"
    },
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is associated with building and orchestrating ETL workflows, which is a primary responsibility of the role.",
      "skill_name": "Azure Data Factory"
    },
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is mentioned in the context of developing and maintaining Lakehouse architecture, indicating its primary relevance to the role.",
      "skill_name": "ADLS"
    },
    {
      "dimension": {
        "display_name": "Programming Languages for Data Work",
        "id": 21,
        "slug": "programming-languages-for-data-work"
      },
      "is_primary": true,
      "layer": "L0",
      "layer_source": "baseline",
      "rationale": "Anchor dimension: Programming Languages for Data Work",
      "skill_name": "SQL"
    },
    {
      "dimension": null,
      "is_primary": true,
      "layer": "L1",
      "layer_source": "llm",
      "rationale": "The skill is tied to building and orchestrating ETL workflows, which is a primary responsibility in the job description.",
      "skill_name": "ETL"
    }
  ],
  "jd_parameters": {
    "certifications": [],
    "clientDetails": null,
    "company": "Our client",
    "companySize": null,
    "ctc": {
      "currency": null,
      "max": null,
      "min": null,
      "period": null,
      "raw": null
    },
    "educationRequirements": [],
    "experience": {
      "max": null,
      "min": 5,
      "raw": "5 years of experience in Data Engineering"
    },
    "industryDomain": "IT Services \u0026 Consulting",
    "knockouts": {
      "certifications": [],
      "ctc": {
        "currency": null,
        "max": null,
        "min": null
      },
      "educationRequirements": [],
      "experience": {
        "max": null,
        "min": 5
      },
      "location": [
        "Bengaluru, India",
        "Hyderabad, India",
        "Gurugram, India"
      ],
      "noticePeriod": null
    },
    "locations": [
      "Bengaluru, India",
      "Hyderabad, India",
      "Gurugram, India"
    ],
    "noticePeriod": null,
    "openToRelocate": false,
    "role": "Data Engineer (Databricks)",
    "roleSynonyms": [
      "Data Engineer",
      "Data Engineer (Azure)",
      "ETL Developer"
    ]
  },
  "jd_role": {
    "display_name": "Data Engineer (Databricks)",
    "rationale": null,
    "role_aliases": [
      "Data Engineer",
      "Data Engineer (Azure)",
      "ETL Developer"
    ],
    "role_archetype": "Data",
    "slug": ""
  },
  "nano_parsed": {
    "JD_type": "pass",
    "about_company": {
      "source_marker": {
        "first_5_words": "Our client is a globally",
        "last_5_words": "organizations worldwide."
      },
      "text": "Our client is a globally recognized AI \u0026\n\u2022 Analytics company known for delivering cutting-edge data solutions, advanced analytics, data engineering, governance, MLOps, and Generative AI innovations to leading organizations worldwide.",
      "word_count": 40
    },
    "ai_kras": [],
    "certifications": [],
    "company_name": "Our client",
    "ctc": null,
    "domain": {
      "primary": {
        "aliases": [
          "ITES",
          "BPO"
        ],
        "domain": "IT Services \u0026 Consulting"
      },
      "secondary": null
    },
    "education": [],
    "experience": {
      "max": null,
      "min": 5,
      "raw": "5 years of experience in Data Engineering"
    },
    "job_locations": [
      {
        "aliases": [
          "Bangalore"
        ],
        "city": "Bengaluru",
        "country": "India",
        "state": "Karnataka",
        "work_mode": "null"
      },
      {
        "aliases": [],
        "city": "Hyderabad",
        "country": "India",
        "state": "Telangana",
        "work_mode": "null"
      },
      {
        "aliases": [
          "Gurgaon"
        ],
        "city": "Gurugram",
        "country": "India",
        "state": "Haryana",
        "work_mode": "null"
      }
    ],
    "role": "Data Engineer (Databricks)",
    "role_aliases": [
      {
        "name": "Data Engineer",
        "reasoning": "generalized form of the picked role",
        "relation": "synonym"
      },
      {
        "name": "Data Engineer (Azure)",
        "reasoning": "specific focus on Azure technologies",
        "relation": "synonym"
      },
      {
        "name": "ETL Developer",
        "reasoning": "role involves significant ETL development",
        "relation": "synonym"
      }
    ],
    "role_archetype": "Data",
    "roles_and_responsibilities": [
      {
        "bullet_count": 6,
        "heading": "Responsibilities",
        "heading_was_present": true,
        "source_marker": {
          "first_5_words": "1. Design and develop scalable",
          "last_5_words": "and DevOps teams on enterprise"
        },
        "text": "1. Design and develop scalable data pipelines using PySpark \u0026\n\u2022 Spark SQL on Azure Databricks 2. Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3. Develop and maintain Lakehouse architecture using ADLS \u0026\n\u2022 Databricks 4. Perform data cleansing, normalization, deduplication, and transformations 5. Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6. Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives",
        "word_count": 66
      },
      {
        "bullet_count": 6,
        "heading": "Requirements",
        "heading_was_present": true,
        "source_marker": {
          "first_5_words": "1. 5 years of experience",
          "last_5_words": "short notice candidates preferred"
        },
        "text": "1. 5 years of experience in Data Engineering 2. Strong hands-on experience with Azure Databricks, Azure Data Factory, ADLS 3. Expertise in PySpark, Spark SQL, SQL, ETL Development 4. Good understanding of data architecture and cloud platforms 5. Robust communication and stakeholder management skills 6. Immediate to short notice candidates preferred",
        "word_count": 66
      }
    ],
    "urls": []
  },
  "rejected": false,
  "rejection_reason": null,
  "role": {
    "anchor_type": null,
    "display_name": "Data Engineer",
    "id": 2,
    "resolution": "in_db",
    "slug": "data-engineer"
  },
  "run_id": "c9eddb74-c671-4d76-b490-b9d9b04cd64e",
  "secondary_skills": [],
  "skill_layers": [
    {
      "label": "Anchor",
      "layer": "L0",
      "skills": [
        {
          "dimension": {
            "display_name": "Programming Languages for Data Work",
            "id": 21,
            "slug": "programming-languages-for-data-work"
          },
          "name": "SQL",
          "rationale": "Anchor dimension: Programming Languages for Data Work"
        }
      ]
    },
    {
      "label": "Primary",
      "layer": "L1",
      "skills": [
        {
          "name": "PySpark",
          "rationale": "The skill is backed by a responsibility that involves designing and developing scalable data pipelines using PySpark."
        },
        {
          "name": "Spark SQL",
          "rationale": "The skill is mentioned in the context of building and orchestrating ETL workflows, indicating its primary importance in the role."
        },
        {
          "name": "Azure Databricks",
          "rationale": "The skill is referenced in relation to Spark SQL and ETL workflows, highlighting its primary role in the data engineering tasks."
        },
        {
          "name": "Azure Data Factory",
          "rationale": "The skill is associated with building and orchestrating ETL workflows, which is a primary responsibility of the role."
        },
        {
          "name": "ADLS",
          "rationale": "The skill is mentioned in the context of developing and maintaining Lakehouse architecture, indicating its primary relevance to the role."
        },
        {
          "name": "ETL",
          "rationale": "The skill is tied to building and orchestrating ETL workflows, which is a primary responsibility in the job description."
        }
      ]
    }
  ],
  "stage3_signals": {
    "alias_found": true,
    "alias_match_roles": [
      {
        "display_name": "Data Engineer",
        "kra_matches": null,
        "matched_count": null,
        "matched_skills": null,
        "role_id": 2,
        "score": 1.0,
        "slug": "data-engineer",
        "total_count": null
      }
    ],
    "kra_match_roles": [
      {
        "display_name": "Data Engineer",
        "kra_matches": [
          {
            "kra_text": "Implements data transformation, cleansing, deduplication, and enrichment logic to convert raw source data into analytics-ready curated datasets.",
            "sentence": "Perform data cleansing, normalization, deduplication, and transformations 5.",
            "similarity": 0.6798
          },
          {
            "kra_text": "Develops batch and real-time streaming data pipelines using Apache Spark, Apache Kafka, Apache Flink, or Airflow for data movement and processing at scale.",
            "sentence": "Design and develop scalable data pipelines using PySpark \u0026",
            "similarity": 0.6795
          },
          {
            "kra_text": "Monitors pipeline health, SLA breach alerts, and job failure notifications, and performs root cause analysis for data pipeline incidents.",
            "sentence": "Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6.",
            "similarity": 0.6341
          }
        ],
        "matched_count": null,
        "matched_skills": null,
        "role_id": 2,
        "score": 0.6645,
        "slug": "data-engineer",
        "total_count": null
      },
      {
        "display_name": "DevOps Engineer",
        "kra_matches": [
          {
            "kra_text": "Monitors CI/CD pipeline reliability, identifies bottlenecks in delivery workflows, and improves deployment frequency, lead time, and failure recovery rate.",
            "sentence": "Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6.",
            "similarity": 0.6039
          },
          {
            "kra_text": "Collaborates with development teams to improve build processes, reduce deployment friction, containerize applications, and adopt DevOps best practices.",
            "sentence": "Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives",
            "similarity": 0.5285
          },
          {
            "kra_text": "Collaborates with development teams to improve build processes, reduce deployment friction, containerize applications, and adopt DevOps best practices.",
            "sentence": "Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3.",
            "similarity": 0.4011
          }
        ],
        "matched_count": null,
        "matched_skills": null,
        "role_id": 10,
        "score": 0.5112,
        "slug": "devops-engineer",
        "total_count": null
      },
      {
        "display_name": "ML Engineer",
        "kra_matches": [
          {
            "kra_text": "Monitors production model behavior for data drift, concept drift, and prediction performance degradation using monitoring dashboards and alerting.",
            "sentence": "Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6.",
            "similarity": 0.5103
          },
          {
            "kra_text": "Designs end-to-end ML training pipelines and model inference workflows using TensorFlow, PyTorch, or scikit-learn on cloud ML platforms.",
            "sentence": "Design and develop scalable data pipelines using PySpark \u0026",
            "similarity": 0.4887
          },
          {
            "kra_text": "Prepares, cleans, and transforms training datasets, manages feature stores, and builds feature engineering pipelines for model training.",
            "sentence": "Perform data cleansing, normalization, deduplication, and transformations 5.",
            "similarity": 0.4758
          }
        ],
        "matched_count": null,
        "matched_skills": null,
        "role_id": 3,
        "score": 0.4916,
        "slug": "ml-engineer",
        "total_count": null
      },
      {
        "display_name": "MLOps Engineer",
        "kra_matches": [
          {
            "kra_text": "Supports ML platform incidents by diagnosing model serving failures, feature store pipeline breaks, and training environment configuration issues.",
            "sentence": "Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6.",
            "similarity": 0.52
          },
          {
            "kra_text": "Validates model performance benchmarks, data schema contracts, and system integration health before signing off on production release readiness.",
            "sentence": "Perform data cleansing, normalization, deduplication, and transformations 5.",
            "similarity": 0.4533
          },
          {
            "kra_text": "Automates ML platform operations including scheduled retraining triggers, pipeline orchestration, evaluation workflows, and alerting configuration.",
            "sentence": "Build and orchestrate ETL workflows using Azure Data Factory (ADF) 3.",
            "similarity": 0.4335
          }
        ],
        "matched_count": null,
        "matched_skills": null,
        "role_id": 16,
        "score": 0.4689,
        "slug": "ml-ops-engineer",
        "total_count": null
      },
      {
        "display_name": "AI Engineer",
        "kra_matches": [
          {
            "kra_text": "Monitors AI feature behavior in production including response quality metrics, latency percentiles, token cost per request, and error rates.",
            "sentence": "Monitor data pipelines, troubleshoot issues, and ensure smooth production support 6.",
            "similarity": 0.4684
          },
          {
            "kra_text": "Designs and implements prompt engineering workflows, few-shot examples, chain-of-thought patterns, and structured output parsing for AI feature pipelines.",
            "sentence": "Design and develop scalable data pipelines using PySpark \u0026",
            "similarity": 0.4609
          },
          {
            "kra_text": "Integrates AI model API responses with application business logic, database writes, event publishing, and downstream service orchestration.",
            "sentence": "Collaborate with BI, Data Science, and DevOps teams on enterprise analytics initiatives",
            "similarity": 0.4268
          }
        ],
        "matched_count": null,
        "matched_skills": null,
        "role_id": 13,
        "score": 0.452,
        "slug": "ai-engineer",
        "total_count": null
      }
    ],
    "skill_match_roles": [
      {
        "display_name": "Fullstack Developer",
        "kra_matches": null,
        "matched_count": 1,
        "matched_skills": [
          "SQL"
        ],
        "role_id": 15,
        "score": 0.1429,
        "slug": "full-stack-engineer",
        "total_count": 7
      },
      {
        "display_name": "Backend Developer",
        "kra_matches": null,
        "matched_count": 1,
        "matched_skills": [
          "SQL"
        ],
        "role_id": 1,
        "score": 0.1429,
        "slug": "backend-engineer",
        "total_count": 7
      },
      {
        "display_name": "Data Engineer",
        "kra_matches": null,
        "matched_count": 1,
        "matched_skills": [
          "SQL"
        ],
        "role_id": 2,
        "score": 0.1429,
        "slug": "data-engineer",
        "total_count": 7
      },
      {
        "display_name": "Pega Developer",
        "kra_matches": null,
        "matched_count": 1,
        "matched_skills": [
          "SQL"
        ],
        "role_id": 24,
        "score": 0.1429,
        "slug": "pega-developer",
        "total_count": 7
      },
      {
        "display_name": "Engineering Manager",
        "kra_matches": null,
        "matched_count": 1,
        "matched_skills": [
          "SQL"
        ],
        "role_id": 121,
        "score": 0.1429,
        "slug": "engineering-manager",
        "total_count": 7
      }
    ]
  },
  "stage4_decision": {
    "alias_collision_detected": false,
    "case": "A",
    "chosen_role": {
      "display_name": "Data Engineer",
      "kra_matches": null,
      "matched_count": null,
      "matched_skills": null,
      "role_id": 2,
      "score": 1.0,
      "slug": "data-engineer",
      "total_count": null
    },
    "confidence": 1.0,
    "consensus_evidence": {},
    "is_new_role": false,
    "llm2_fired": false,
    "llm2_reasoning": null,
    "matched_dimensions": [],
    "matched_kras": [],
    "matched_skills": [],
    "new_role_display_name": null,
    "new_role_domain": null,
    "new_role_slug": null,
    "queued": false,
    "reasoning": "Exact alias hit on data-engineer (1.0) \u2014 no other alias at this confidence; skill_top full-stack-engineer 0.14 and KRA_top data-engineer 0.66 do not contradict",
    "role_aliases": [],
    "sub_role": null
  },
  "stage5_updates": {
    "centroid_n_after": 526,
    "centroid_updated": true,
    "collision_log_id": null,
    "new_kra_attached": null,
    "new_skills_attached": [
      {
        "is_primary": true,
        "queue_id": 26264,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "PySpark",
        "status": "pending"
      },
      {
        "is_primary": true,
        "queue_id": 26265,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "Spark SQL",
        "status": "pending"
      },
      {
        "is_primary": true,
        "queue_id": 26266,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "Azure Databricks",
        "status": "pending"
      },
      {
        "is_primary": true,
        "queue_id": 26267,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "Azure Data Factory",
        "status": "pending"
      },
      {
        "is_primary": true,
        "queue_id": 26268,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "ADLS",
        "status": "pending"
      },
      {
        "is_primary": true,
        "queue_id": 26269,
        "role_display_name": "Data Engineer",
        "role_slug": "data-engineer",
        "skill_name": "ETL",
        "status": "pending"
      }
    ],
    "queue_entry_id": null,
    "v3_pipeline_triggered": false,
    "v3_role_slug": null,
    "v3_run_id": null
  }
}
API 2 — extract-details
{
  "alias_matches": [
    {
      "alias_persist_skipped_reason": "TODO: REMOVE AFTER TESTING \u2014 alias DB write disabled",
      "alias_persisted": false,
      "existing_alias_id": 2004,
      "existing_alias_text": "Apache Spark",
      "input_term": "PySpark",
      "matched_canonical": {
        "category_id": 5,
        "display_name": "Apache Spark",
        "id": 1350,
        "is_also_category": false,
        "is_extractable": true,
        "skill_nature": "FRAMEWORK",
        "slug": "apache-spark",
        "sub_category_id": 1021,
        "typical_lifespan": "EVERGREEN",
        "volatility": "STABLE"
      },
      "matched_via": "embedding_alias"
    },
    {
      "alias_persist_skipped_reason": "alias_text already exists for this canonical skill",
      "alias_persisted": false,
      "existing_alias_id": 271,
      "existing_alias_text": "SQL",
      "input_term": "SQL",
      "matched_canonical": {
        "category_id": 6,
        "display_name": "SQL",
        "id": 101,
        "is_also_category": false,
        "is_extractable": true,
        "skill_nature": "LANGUAGE",
        "slug": "sql",
        "sub_category_id": 97,
        "typical_lifespan": "EVERGREEN",
        "volatility": "STABLE"
      },
      "matched_via": "alias"
    }
  ],
  "candidate_roles": [
    {
      "display_name": "Data Engineer",
      "id": 2,
      "rationale": null,
      "role_archetype": null,
      "slug": "data-engineer",
      "source": "db"
    },
    {
      "display_name": "Pega Developer",
      "id": 24,
      "rationale": null,
      "role_archetype": null,
      "slug": "pega-developer",
      "source": "db"
    },
    {
      "display_name": "Backend Developer",
      "id": 1,
      "rationale": null,
      "role_archetype": "A Backend Engineer designs, builds, and maintains the server-side logic and data handling that power applications and services. They focus on implementing reliable business functionality, integrating with other systems, and ensuring the backend is scalable, maintainable, and observable.",
      "slug": "backend-engineer",
      "source": "db"
    },
    {
      "display_name": "Fullstack Developer",
      "id": 15,
      "rationale": null,
      "role_archetype": null,
      "slug": "full-stack-engineer",
      "source": "db"
    },
    {
      "display_name": "Fullstack Developer",
      "id": 435,
      "rationale": null,
      "role_archetype": "Engineering",
      "slug": "fullstack-developer",
      "source": "db"
    },
    {
      "display_name": "WordPress Dev",
      "id": 227,
      "rationale": null,
      "role_archetype": "Engineering",
      "slug": "wordpress-dev",
      "source": "db"
    },
    {
      "display_name": "Engineering Manager",
      "id": 121,
      "rationale": null,
      "role_archetype": null,
      "slug": "engineering-manager",
      "source": "db"
    }
  ],
  "chosen_role": {
    "display_name": "Data Engineer",
    "id": 2,
    "rationale": "Exact alias hit on data-engineer (1.0) \u2014 no other alias at this confidence; skill_top full-stack-engineer 0.14 and KRA_top data-engineer 0.66 do not contradict",
    "role_archetype": null,
    "slug": "data-engineer",
    "source": "db"
  },
  "dimensions": [
    {
      "dimension": {
        "difficulty_hint": "well_known",
        "display_name": "ETL and ELT Tooling",
        "id": 24,
        "rationale": "Packaged tools for extracting, loading, and transforming data across systems. This dimension covers connector-based ingestion, transformation frameworks, and managed integration products.",
        "slug": "etl-and-elt-tooling",
        "source": "db"
      },
      "input_skill": "PySpark",
      "llm_role": null,
      "roles_from_db": [
        {
          "display_name": "Data Engineer",
          "id": 2,
          "rationale": null,
          "role_archetype": null,
          "slug": "data-engineer",
          "source": "db"
        }
      ]
    },
    {
      "dimension": {
        "difficulty_hint": "well_known",
        "display_name": "Pega Programming Languages \u0026 DSLs",
        "id": 267,
        "rationale": "Programming languages and domain-specific languages used in Pega development.",
        "slug": "pega-programming-languages-dsls",
        "source": "db"
      },
      "input_skill": "SQL",
      "llm_role": null,
      "roles_from_db": [
        {
          "display_name": "Pega Developer",
          "id": 24,
          "rationale": null,
          "role_archetype": null,
          "slug": "pega-developer",
          "source": "db"
        }
      ]
    },
    {
      "dimension": {
        "difficulty_hint": "well_known",
        "display_name": "Programming Languages",
        "id": 1,
        "rationale": "Core programming languages and markup used in WordPress development.",
        "slug": "programming-languages",
        "source": "db"
      },
      "input_skill": "SQL",
      "llm_role": null,
      "roles_from_db": [
        {
          "display_name": "Backend Developer",
          "id": 1,
          "rationale": null,
          "role_archetype": "A Backend Engineer designs, builds, and maintains the server-side logic and data handling that power applications and services. They focus on implementing reliable business functionality, integrating with other systems, and ensuring the backend is scalable, maintainable, and observable.",
          "slug": "backend-engineer",
          "source": "db"
        },
        {
          "display_name": "Fullstack Developer",
          "id": 15,
          "rationale": null,
          "role_archetype": null,
          "slug": "full-stack-engineer",
          "source": "db"
        },
        {
          "display_name": "Fullstack Developer",
          "id": 435,
          "rationale": null,
          "role_archetype": "Engineering",
          "slug": "fullstack-developer",
          "source": "db"
        },
        {
          "display_name": "WordPress Dev",
          "id": 227,
          "rationale": null,
          "role_archetype": "Engineering",
          "slug": "wordpress-dev",
          "source": "db"
        }
      ]
    },
    {
      "dimension": {
        "difficulty_hint": "well_known",
        "display_name": "Programming Languages \u0026 DSLs",
        "id": 475,
        "rationale": "Oversee and guide the selection and effective use of programming and domain\u2010specific languages in software projects.",
        "slug": "programming-languages-dsls",
        "source": "db"
      },
      "input_skill": "SQL",
      "llm_role": null,
      "roles_from_db": [
        {
          "display_name": "Engineering Manager",
          "id": 121,
          "rationale": null,
          "role_archetype": null,
          "slug": "engineering-manager",
          "source": "db"
        }
      ]
    },
    {
      "dimension": {
        "difficulty_hint": "well_known",
        "display_name": "Programming Languages for Data Work",
        "id": 21,
        "rationale": "Languages used to implement data pipelines, transformations, and operational glue. This is the primary coding surface for building ingestion, enrichment, and automation logic in data engineering.",
        "slug": "programming-languages-for-data-work",
        "source": "db"
      },
      "input_skill": "SQL",
      "llm_role": null,
      "roles_from_db": [
        {
          "display_name": "Data Engineer",
          "id": 2,
          "rationale": null,
          "role_archetype": null,
          "slug": "data-engineer",
          "source": "db"
        }
      ]
    }
  ],
  "input_final_skills": [
    "PySpark",
    "Spark SQL",
    "Azure Databricks",
    "Azure Data Factory",
    "ADLS",
    "SQL",
    "ETL"
  ],
  "input_llm_skills": [
    "PySpark",
    "Spark SQL",
    "Azure Databricks",
    "Azure Data Factory",
    "ADLS",
    "SQL",
    "ETL"
  ],
  "new_aliases_persisted": 0,
  "run_id": "c9eddb74-c671-4d76-b490-b9d9b04cd64e",
  "skills_detail": [
    {
      "aliases_in_db": [
        {
          "alias_text": "Apache Spark",
          "alias_type": "CANONICAL",
          "id": 2004,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        },
        {
          "alias_text": "apache spark 3",
          "alias_type": "VERSION",
          "id": 2006,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        },
        {
          "alias_text": "spark",
          "alias_type": "VERSION",
          "id": 2510,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        },
        {
          "alias_text": "spark 3",
          "alias_type": "VERSION",
          "id": 2007,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        },
        {
          "alias_text": "spark 3.x",
          "alias_type": "VERSION",
          "id": 2009,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        },
        {
          "alias_text": "spark3",
          "alias_type": "VERSION",
          "id": 2008,
          "is_primary": false,
          "match_strategy": "CASE_INSENSITIVE"
        }
      ],
      "canonical": {
        "category_id": 5,
        "display_name": "Apache Spark",
        "id": 1350,
        "is_also_category": false,
        "is_extractable": true,
        "skill_nature": "FRAMEWORK",
        "slug": "apache-spark",
        "sub_category_id": 1021,
        "typical_lifespan": "EVERGREEN",
        "volatility": "STABLE"
      },
      "dimensions": [
        {
          "dimension": {
            "difficulty_hint": "well_known",
            "display_name": "ETL and ELT Tooling",
            "id": 24,
            "rationale": "Packaged tools for extracting, loading, and transforming data across systems. This dimension covers connector-based ingestion, transformation frameworks, and managed integration products.",
            "slug": "etl-and-elt-tooling",
            "source": "db"
          },
          "input_skill": "PySpark",
          "llm_role": null,
          "roles_from_db": [
            {
              "display_name": "Data Engineer",
              "id": 2,
              "rationale": null,
              "role_archetype": null,
              "slug": "data-engineer",
              "source": "db"
            }
          ]
        }
      ],
      "input_skill": "PySpark",
      "matched_via": "embedding_alias",
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": null,
      "source_tag": "db",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [],
      "canonical": null,
      "dimensions": [],
      "input_skill": "Spark SQL",
      "matched_via": null,
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": {
        "derived": {
          "category": "Data Engineering Tools",
          "skill_nature": "LANGUAGE",
          "sub_category": "general",
          "typical_lifespan": "MULTI_YEAR",
          "version_strategy": "UNVERSIONED",
          "volatility": "MEDIUM"
        },
        "enrichment": null,
        "keep_log": [],
        "locked_dimensions": [],
        "merge_log": [],
        "placed": null,
        "relationships": null,
        "skill_id": "spark-sql",
        "split_log": [],
        "typed": null,
        "warnings": []
      },
      "source_tag": "llm",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [],
      "canonical": null,
      "dimensions": [],
      "input_skill": "Azure Databricks",
      "matched_via": null,
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": {
        "derived": {
          "category": "Cloud Platforms",
          "skill_nature": "PLATFORM",
          "sub_category": "general",
          "typical_lifespan": "MULTI_YEAR",
          "version_strategy": "UNVERSIONED",
          "volatility": "MEDIUM"
        },
        "enrichment": null,
        "keep_log": [],
        "locked_dimensions": [],
        "merge_log": [],
        "placed": null,
        "relationships": null,
        "skill_id": "azure-databricks",
        "split_log": [],
        "typed": null,
        "warnings": []
      },
      "source_tag": "llm",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [],
      "canonical": null,
      "dimensions": [],
      "input_skill": "Azure Data Factory",
      "matched_via": null,
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": {
        "derived": {
          "category": "Cloud Platforms",
          "skill_nature": "PLATFORM",
          "sub_category": "general",
          "typical_lifespan": "MULTI_YEAR",
          "version_strategy": "UNVERSIONED",
          "volatility": "MEDIUM"
        },
        "enrichment": null,
        "keep_log": [],
        "locked_dimensions": [],
        "merge_log": [],
        "placed": null,
        "relationships": null,
        "skill_id": "azure-data-factory",
        "split_log": [],
        "typed": null,
        "warnings": []
      },
      "source_tag": "llm",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [],
      "canonical": null,
      "dimensions": [],
      "input_skill": "ADLS",
      "matched_via": null,
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": {
        "derived": {
          "category": "Cloud Platforms",
          "skill_nature": "PLATFORM",
          "sub_category": "general",
          "typical_lifespan": "MULTI_YEAR",
          "version_strategy": "UNVERSIONED",
          "volatility": "MEDIUM"
        },
        "enrichment": null,
        "keep_log": [],
        "locked_dimensions": [],
        "merge_log": [],
        "placed": null,
        "relationships": null,
        "skill_id": "adls",
        "split_log": [],
        "typed": null,
        "warnings": []
      },
      "source_tag": "llm",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [
        {
          "alias_text": "SQL",
          "alias_type": "CANONICAL",
          "id": 271,
          "is_primary": true,
          "match_strategy": "CASE_INSENSITIVE"
        }
      ],
      "canonical": {
        "category_id": 6,
        "display_name": "SQL",
        "id": 101,
        "is_also_category": false,
        "is_extractable": true,
        "skill_nature": "LANGUAGE",
        "slug": "sql",
        "sub_category_id": 97,
        "typical_lifespan": "EVERGREEN",
        "volatility": "STABLE"
      },
      "dimensions": [
        {
          "dimension": {
            "difficulty_hint": "well_known",
            "display_name": "Pega Programming Languages \u0026 DSLs",
            "id": 267,
            "rationale": "Programming languages and domain-specific languages used in Pega development.",
            "slug": "pega-programming-languages-dsls",
            "source": "db"
          },
          "input_skill": "SQL",
          "llm_role": null,
          "roles_from_db": [
            {
              "display_name": "Pega Developer",
              "id": 24,
              "rationale": null,
              "role_archetype": null,
              "slug": "pega-developer",
              "source": "db"
            }
          ]
        },
        {
          "dimension": {
            "difficulty_hint": "well_known",
            "display_name": "Programming Languages",
            "id": 1,
            "rationale": "Core programming languages and markup used in WordPress development.",
            "slug": "programming-languages",
            "source": "db"
          },
          "input_skill": "SQL",
          "llm_role": null,
          "roles_from_db": [
            {
              "display_name": "Backend Developer",
              "id": 1,
              "rationale": null,
              "role_archetype": "A Backend Engineer designs, builds, and maintains the server-side logic and data handling that power applications and services. They focus on implementing reliable business functionality, integrating with other systems, and ensuring the backend is scalable, maintainable, and observable.",
              "slug": "backend-engineer",
              "source": "db"
            },
            {
              "display_name": "Fullstack Developer",
              "id": 15,
              "rationale": null,
              "role_archetype": null,
              "slug": "full-stack-engineer",
              "source": "db"
            },
            {
              "display_name": "Fullstack Developer",
              "id": 435,
              "rationale": null,
              "role_archetype": "Engineering",
              "slug": "fullstack-developer",
              "source": "db"
            },
            {
              "display_name": "WordPress Dev",
              "id": 227,
              "rationale": null,
              "role_archetype": "Engineering",
              "slug": "wordpress-dev",
              "source": "db"
            }
          ]
        },
        {
          "dimension": {
            "difficulty_hint": "well_known",
            "display_name": "Programming Languages \u0026 DSLs",
            "id": 475,
            "rationale": "Oversee and guide the selection and effective use of programming and domain\u2010specific languages in software projects.",
            "slug": "programming-languages-dsls",
            "source": "db"
          },
          "input_skill": "SQL",
          "llm_role": null,
          "roles_from_db": [
            {
              "display_name": "Engineering Manager",
              "id": 121,
              "rationale": null,
              "role_archetype": null,
              "slug": "engineering-manager",
              "source": "db"
            }
          ]
        },
        {
          "dimension": {
            "difficulty_hint": "well_known",
            "display_name": "Programming Languages for Data Work",
            "id": 21,
            "rationale": "Languages used to implement data pipelines, transformations, and operational glue. This is the primary coding surface for building ingestion, enrichment, and automation logic in data engineering.",
            "slug": "programming-languages-for-data-work",
            "source": "db"
          },
          "input_skill": "SQL",
          "llm_role": null,
          "roles_from_db": [
            {
              "display_name": "Data Engineer",
              "id": 2,
              "rationale": null,
              "role_archetype": null,
              "slug": "data-engineer",
              "source": "db"
            }
          ]
        }
      ],
      "input_skill": "SQL",
      "matched_via": "alias",
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": null,
      "source_tag": "db",
      "was_in_llm_skills": true
    },
    {
      "aliases_in_db": [],
      "canonical": null,
      "dimensions": [],
      "input_skill": "ETL",
      "matched_via": null,
      "new_alias_persisted": false,
      "new_alias_text": null,
      "new_skill_meta": {
        "derived": {
          "category": "Data Engineering Tools",
          "skill_nature": "PRACTICE",
          "sub_category": "general",
          "typical_lifespan": "MULTI_YEAR",
          "version_strategy": "UNVERSIONED",
          "volatility": "MEDIUM"
        },
        "enrichment": null,
        "keep_log": [],
        "locked_dimensions": [],
        "merge_log": [],
        "placed": null,
        "relationships": null,
        "skill_id": "etl",
        "split_log": [],
        "typed": null,
        "warnings": []
      },
      "source_tag": "llm",
      "was_in_llm_skills": true
    }
  ],
  "unmatched_skills": [
    "Spark SQL",
    "Azure Databricks",
    "Azure Data Factory",
    "ADLS",
    "ETL"
  ]
}
API 3 — final-role-output
{
  "chosen_role": {
    "display_name": "Data Engineer",
    "id": 2,
    "rationale": "Exact alias hit on data-engineer (1.0) \u2014 no other alias at this confidence; skill_top full-stack-engineer 0.14 and KRA_top data-engineer 0.66 do not contradict",
    "role_archetype": null,
    "slug": "data-engineer",
    "source": "db"
  },
  "chosen_role_resolution": "in_db",
  "final_input_skills": [
    {
      "skill": "PySpark",
      "tag": "in_db"
    },
    {
      "skill": "Spark SQL",
      "tag": "new"
    },
    {
      "skill": "Azure Databricks",
      "tag": "new"
    },
    {
      "skill": "Azure Data Factory",
      "tag": "new"
    },
    {
      "skill": "ADLS",
      "tag": "new"
    },
    {
      "skill": "SQL",
      "tag": "in_db"
    },
    {
      "skill": "ETL",
      "tag": "new"
    }
  ],
  "llm_cost_api1_usd": null,
  "llm_cost_api2_usd": null,
  "llm_cost_api3_usd": null,
  "llm_cost_total_usd": null,
  "persistence": {
    "items": [
      {
        "chosen_role_id": 2,
        "dimension": {
          "difficulty_hint": "well_known",
          "display_name": "ETL and ELT Tooling",
          "id": 24,
          "rationale": "Packaged tools for extracting, loading, and transforming data across systems. This dimension covers connector-based ingestion, transformation frameworks, and managed integration products.",
          "slug": "etl-and-elt-tooling",
          "source": "db"
        },
        "dimension_id": 24,
        "input_skill": "PySpark",
        "llm_role": null,
        "matched_chosen_role": true,
        "outcome_line": "Skipped \u2014 no persistable v3 meta for new skill",
        "role_dimension_saved": false,
        "roles_from_db": [
          {
            "display_name": "Data Engineer",
            "id": 2,
            "rationale": null,
            "role_archetype": null,
            "slug": "data-engineer",
            "source": "db"
          }
        ],
        "skill_dimension_saved": false,
        "skill_id": null,
        "skill_tag": "new",
        "skipped_reason": "skill_not_in_db_v3_proposed"
      },
      {
        "chosen_role_id": 2,
        "dimension": {
          "difficulty_hint": "well_known",
          "display_name": "Pega Programming Languages \u0026 DSLs",
          "id": 267,
          "rationale": "Programming languages and domain-specific languages used in Pega development.",
          "slug": "pega-programming-languages-dsls",
          "source": "db"
        },
        "dimension_id": 267,
        "input_skill": "SQL",
        "llm_role": null,
        "matched_chosen_role": false,
        "outcome_line": "Existing dimension (library) \u00b7 Role\u2194dimension skipped (dimension not under chosen role)",
        "role_dimension_saved": false,
        "roles_from_db": [
          {
            "display_name": "Pega Developer",
            "id": 24,
            "rationale": null,
            "role_archetype": null,
            "slug": "pega-developer",
            "source": "db"
          }
        ],
        "skill_dimension_saved": true,
        "skill_id": 101,
        "skill_tag": "in_db",
        "skipped_reason": null
      },
      {
        "chosen_role_id": 2,
        "dimension": {
          "difficulty_hint": "well_known",
          "display_name": "Programming Languages",
          "id": 1,
          "rationale": "Core programming languages and markup used in WordPress development.",
          "slug": "programming-languages",
          "source": "db"
        },
        "dimension_id": 1,
        "input_skill": "SQL",
        "llm_role": null,
        "matched_chosen_role": false,
        "outcome_line": "Existing dimension (library) \u00b7 Role\u2194dimension skipped (dimension not under chosen role)",
        "role_dimension_saved": false,
        "roles_from_db": [
          {
            "display_name": "Backend Developer",
            "id": 1,
            "rationale": null,
            "role_archetype": "A Backend Engineer designs, builds, and maintains the server-side logic and data handling that power applications and services. They focus on implementing reliable business functionality, integrating with other systems, and ensuring the backend is scalable, maintainable, and observable.",
            "slug": "backend-engineer",
            "source": "db"
          },
          {
            "display_name": "Fullstack Developer",
            "id": 15,
            "rationale": null,
            "role_archetype": null,
            "slug": "full-stack-engineer",
            "source": "db"
          },
          {
            "display_name": "Fullstack Developer",
            "id": 435,
            "rationale": null,
            "role_archetype": "Engineering",
            "slug": "fullstack-developer",
            "source": "db"
          },
          {
            "display_name": "WordPress Dev",
            "id": 227,
            "rationale": null,
            "role_archetype": "Engineering",
            "slug": "wordpress-dev",
            "source": "db"
          }
        ],
        "skill_dimension_saved": true,
        "skill_id": 101,
        "skill_tag": "in_db",
        "skipped_reason": null
      },
      {
        "chosen_role_id": 2,
        "dimension": {
          "difficulty_hint": "well_known",
          "display_name": "Programming Languages \u0026 DSLs",
          "id": 475,
          "rationale": "Oversee and guide the selection and effective use of programming and domain\u2010specific languages in software projects.",
          "slug": "programming-languages-dsls",
          "source": "db"
        },
        "dimension_id": 475,
        "input_skill": "SQL",
        "llm_role": null,
        "matched_chosen_role": false,
        "outcome_line": "Existing dimension (library) \u00b7 Role\u2194dimension skipped (dimension not under chosen role)",
        "role_dimension_saved": false,
        "roles_from_db": [
          {
            "display_name": "Engineering Manager",
            "id": 121,
            "rationale": null,
            "role_archetype": null,
            "slug": "engineering-manager",
            "source": "db"
          }
        ],
        "skill_dimension_saved": true,
        "skill_id": 101,
        "skill_tag": "in_db",
        "skipped_reason": null
      },
      {
        "chosen_role_id": 2,
        "dimension": {
          "difficulty_hint": "well_known",
          "display_name": "Programming Languages for Data Work",
          "id": 21,
          "rationale": "Languages used to implement data pipelines, transformations, and operational glue. This is the primary coding surface for building ingestion, enrichment, and automation logic in data engineering.",
          "slug": "programming-languages-for-data-work",
          "source": "db"
        },
        "dimension_id": 21,
        "input_skill": "SQL",
        "llm_role": null,
        "matched_chosen_role": true,
        "outcome_line": "Existing dimension (library) \u00b7 Role\u2194dimension saved",
        "role_dimension_saved": true,
        "roles_from_db": [
          {
            "display_name": "Data Engineer",
            "id": 2,
            "rationale": null,
            "role_archetype": null,
            "slug": "data-engineer",
            "source": "db"
          }
        ],
        "skill_dimension_saved": true,
        "skill_id": 101,
        "skill_tag": "in_db",
        "skipped_reason": null
      }
    ],
    "new_skills_created": 0,
    "role_dimension_saved": 0,
    "skill_dimension_saved": 0,
    "skipped": 1
  },
  "planner_output": null,
  "run_id": "c9eddb74-c671-4d76-b490-b9d9b04cd64e"
}

LLM Calls

Every model call made for this run, in pipeline order. Click a card to see the model's response.

Loading…