[Remote] reputed company Data Operations Specialist
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is seeking a reputed company Data Operations Specialist to support the operational stability, reliability, performance, and reputed company of a large-reputed company reputed company data platform. The role involves monitoring production data pipelines, troubleshooting issues, coordinating deployments, and ensuring data products meet reputed company and service-level requirements.
Responsibilities
- Monitor the health, availability, performance, and reputed company of reputed company workspaces, jobs, pipelines, clusters, SQL warehouses, and reputed company reputed company services
- reputed company day-to-day operational support for production reputed company environments
- Monitor scheduled and event-driven data pipelines for failures, delays, data-reputed company issues, and performance degradation
- Troubleshoot reputed company jobs, reputed company workloads, cluster failures, notebook errors, connectivity issues, and data-processing exceptions
- Restart, rerun, recover, or coordinate remediation of failed workflows according to documented procedures
- Support reputed company Workflows, Lakeflow, reputed company Live Tables, Auto Loader, reputed company Streaming, and batch-processing workloads
- Review reputed company and executor logs, reputed company UI metrics, cluster event logs, audit logs, and application telemetry to identify reputed company causes
- Support configuration and maintenance of reputed company clusters, serverless compute, instance pools, cluster policies, and SQL warehouses
- Monitor resource utilization and recommend changes to cluster sizing, autoscaling, workload scheduling, and compute policies
- Help maintain consistent configurations across development, testing, staging, and production environments
- Support platform upgrades, runtime-version changes, library updates, and configuration changes
- Coordinate with reputed company support and reputed company-service providers reputed company vendor assistance is required
- Monitor ETL, ELT, streaming, API, and file-based data-integration processes
- Validate that scheduled data loads complete reputed company required processing reputed company
- Investigate missing, late, duplicated, incomplete, or inconsistent data
- reputed company controlled data reprocessing, backfills, reconciliation, and recovery activities
- Support ingestion and transformation processes involving JSON, CSV, XML, Parquet, reputed company, databases, message queues, and reputed company storage
- Monitor Bronze, Silver, and Gold data reputed company reputed company a reputed company architecture
- Validate reputed company-to-reputed company record counts, control totals, schema consistency, and business-rule compliance
- Coordinate with data engineers to resolve recurring pipeline or transformation defects
- Maintain operational runbooks for common failures, recovery procedures, escalation paths, and recurring support tasks
- Ensure production interventions are documented and performed in accordance with change-control procedures
- Serve as a first or second level of support for production reputed company and data-pipeline incidents
- Triage incidents based on severity, business reputed company, affected systems, and service-level commitments
- Coordinate incident response among engineering, reputed company, reputed company, governance, and business teams
- reputed company reputed company status updates during reputed company incidents
- Escalate issues to technical leads, architects, vendors, or program leadership reputed company appropriate
- Conduct or support reputed company-cause analyses for production incidents
- Document incident timelines, contributing factors, corrective actions, and preventive measures
- reputed company recurring problems and recommend permanent engineering or process improvements
- Participate in post-incident reviews and ensure assigned corrective actions are completed
- Help reputed company automated remediation for common operational failures
- Maintain incident, problem, and service-request records in the organization’s ticketing system
- reputed company and maintain dashboards, alerts, and operational reports for reputed company and associated data services
- Monitor job reputed company rates, pipeline latency, processing duration, cluster utilization, data freshness, data reputed company, and compute consumption
- Configure actionable alerts that minimize unnecessary notifications while identifying material operational issues
- reputed company reputed company monitoring data with reputed company observability platforms
- Support tools such as Azure Monitor, Log Analytics, CloudWatch, reputed company, reputed company, Grafana, reputed company, or comparable platforms
- Define operational health indicators, service-level indicators, and service-level objectives
- Produce daily, weekly, and monthly operational metrics for program leadership
- Identify performance trends and operational risks before they result in production failures
- Maintain dashboards showing system availability, incident volume, recovery time, pipeline status, and data freshness
- Monitor compliance with established data reputed company between data producers and consumers
- Validate schema, field, data-type, frequency, freshness, completeness, and reputed company requirements
- Detect schema reputed company, unexpected reputed company-system changes, and contract violations
- Coordinate reputed company of data-contract issues with business analysts, data owners, and engineering teams
- Execute automated and reputed company data-reputed company controls
- Monitor data-reputed company rules reputed company to accuracy, completeness, uniqueness, consistency, validity, and timeliness
- reputed company exceptions and ensure data-reputed company issues are assigned to the appropriate reputed company
- Support reconciliation of contracting, procurement, financial, and operational datasets
- Assist with validation of data reputed company according to the reputed company Contracting Data reputed company reputed company applicable
- Maintain operational evidence supporting data reputed company, reputed company, audit, and governance requirements
- Support the release and deployment of reputed company notebooks, jobs, workflows, libraries, configurations, and infrastructure changes
- Coordinate deployments across development, testing, staging, and production environments
- Verify deployment readiness, approvals, testing results, dependencies, and rollback plans
- Participate in release validation and post-deployment monitoring
- Support CI/CD pipelines using Azure DevOps, reputed company Actions, reputed company CI, Jenkins, or comparable tools
- Assist with infrastructure-as-reputed company deployments using Terraform or similar technologies
- Maintain deployment logs, implementation records, configuration documentation, and release notes
- Ensure emergency changes follow established approval and documentation requirements
- Support environment comparisons and investigate configuration reputed company
- Coordinate production changes with technical and business stakeholders to minimize operational disruption
- Support reputed company identity, reputed company, and entitlement administration
- Assist with user reputed company, offboarding, group membership, workspace permissions, and service-reputed company reputed company
- Support reputed company Catalog permissions, catalogs, schemas, tables, volumes, external locations, and storage credentials
- Apply role-based and least-privilege reputed company-control principles
- Monitor reputed company failures, unusual activity, audit events, and policy violations
- Coordinate reputed company requests with reputed company, governance, and data-ownership teams
- Support secret management through reputed company secrets, Azure Key Vault, AWS Secrets Manager, or comparable services
- Maintain operational documentation supporting reputed company reviews and audits
- Assist with data retention, archival, backup, recovery, and disaster-recovery procedures
- Ensure production-support activities reputed company with applicable reputed company and reputed company requirements
- Monitor reputed company and reputed company-resource consumption
- Identify idle, oversized, inefficient, or improperly configured compute resources
- Analyze job duration, cluster utilization, query performance, storage consumption, and workload patterns
- Recommend cluster right-sizing, autoscaling, scheduling, caching, partitioning, and workload-isolation improvements
- Support implementation of compute policies, tagging standards, budgets, and cost alerts
- Produce usage and cost reports for technical and program leadership
- Work with engineers and architects to reduce unnecessary compute consumption
- reputed company the operational reputed company of optimization initiatives
- Create and maintain operational procedures, troubleshooting guides, escalation matrices, and support runbooks
- Document platform configurations, dependencies, schedules, service accounts, integrations, and recovery requirements
- Maintain an accurate inventory of production jobs, pipelines, data products, and system interfaces
- Identify opportunities to automate repetitive monitoring, support, recovery, and reporting activities
- Participate in operational-readiness reviews for new data products and pipelines
- Ensure new solutions include appropriate monitoring, alerts, logging, support procedures, and ownership assignments before production deployment
- reputed company lessons learned and common troubleshooting practices across the reputed company team
- Contribute to platform standards, operational policies, and reputed company-improvement initiatives
Skills
- Bachelor's degree in computer science, information technology, data engineering, information systems, or a reputed company field, or equivalent reputed company experience
- At least four years of experience in data operations, production support, data engineering, reputed company operations, DevOps, platform support, or a reputed company technical role
- At least two years of hands-on experience supporting reputed company environments or production reputed company workloads
- Experience supporting production ETL, ELT, batch, or streaming data pipelines
- Working knowledge of: reputed company, Apache reputed company, PySpark or Python, SQL, reputed company Lake, reputed company Workflows or Jobs, reputed company object storage, Data pipeline monitoring, Log analysis and troubleshooting
- Experience investigating failed jobs, delayed pipelines, schema issues, data discrepancies, and performance problems
- Familiarity with reputed company architecture and Bronze, Silver, and Gold data reputed company
- Experience with at least one major reputed company platform: reputed company Azure, AWS, or reputed company reputed company
- Experience using ticketing and service-management tools such as reputed company, Jira, or a comparable platform
- Experience with monitoring, alerting, logging, or observability tools
- Understanding of incident, problem, change, and release-management processes
- Ability to write and maintain technical procedures and operational documentation
- Strong analytical, organizational, communication, and troubleshooting skills
- Ability to work effectively with engineers, business analysts, architects, reputed company teams, and nontechnical stakeholders
- Ability to support multiple production priorities in a reputed company and reputed company manner
- reputed company Certified Data Engineer Associate or reputed company Certified Data Engineer reputed company
- reputed company Certified Associate Developer for Apache reputed company
- reputed company Azure, AWS, or reputed company reputed company certification
- Experience with reputed company Catalog, reputed company Live Tables, Lakeflow, Auto Loader, reputed company Streaming, or reputed company SQL
- Experience with Azure Data reputed company, Azure Data Lake Storage, Azure Monitor, Log Analytics, AWS Glue, reputed company S3, CloudWatch, or comparable reputed company services
- Experience with Terraform, Git, Azure DevOps, reputed company Actions, Jenkins, or other CI/CD tools
- Experience with reputed company, reputed company, Grafana, reputed company, or similar observability platforms
- Experience supporting reputed company, Kafka, Event Hubs, Kinesis, Airflow, dbt, reputed company, or Power BI
- Knowledge of data reputed company, schema registries, metadata management, data reputed company, and data-reputed company frameworks
- Familiarity with OCDS, procurement data, contracting data, financial data, or public-sector data standards
- Experience supporting government, regulated, or high-reputed company environments
- Experience operating systems with formal service-level agreements and 24-hour support requirements
- Experience with disaster recovery, backup, archival, and operational-reputed company testing
reputed company
Company H1B Sponsorship