Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Engineer Legion: Director of Production Engineering

Leads DevOps and SRE teams to build reliable, scalable, and secure AWS-based production infrastructure while spending 20-30% time on hands-on architecture and tooling work.

Lead Remote Posted about 8 hours ago We Work Remotely — Programming
What this role involves

Headquarters: Remote, United States

Director of Production Engineering 

Remote, United States

About this Position

Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.

This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.

Responsibilities

  • Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.
  • Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.
  • Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.
  • Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.
  • Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.
  • Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).
  • Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.
  • Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.
  • Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.
  • Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.
  • Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.

Required Qualifications

  • 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.
  • Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).
  • Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues.
  • 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.
  • Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.
  • Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.
  • Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).
  • Proven track record partnering cross-functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.
  • Demonstrated experience leading incident management and on-call practices for high-availability production systems.
  • Bachelor's degree in Computer Science, Engineering, or related field required; Master's degree preferred.

Preferred Qualifications

  • Experience with major cloud providers beyond AWS, such as Google Cloud Platform or Oracle Cloud Infrastructure (OCI).
  • Relevant security certifications (e.g., CISSP, AWS Security Specialty, CKS).
  • Experience with compliance frameworks such as SOC 2, ISO 27001, or HIPAA.
  • 3+ years of experience with Kubernetes or other containerization/orchestration platforms at scale.
  • Experience with Kubernetes-native delivery tooling, including Argo Workflows and Helm.
  • 5+ years of experience leading teams in an agile/scrum environment.
  • Experience building or scaling automated investigation and remediation pipelines for production error classes.

COMPENSATION & BENEFITS

Salary Range: Base Salary Range  $220,000 - $265,000 + Bonus + Stock Equity

At Legion, we offer competitive compensation and benefits packages to all employees. As a fully remote employer, pay for positions is determined using local, national, and industry-specific survey data. 

Our posted salary range is done so in good faith based on national data and may be refined for a candidate's region/town/cost of living. We strive to make competitive offers that allow employees room for future growth. Salaries will be based on the applicant’s location, level of experience, education, and specialized knowledge and skills. Additionally, we consider the external market rate, the amount we have budgeted internally, and the internal equity for the same position within the company. 

Benefits include, but are not limited to:

  • $0 monthly premium and other flexible medical, dental, and vision plans effective on the first day of employment
  • 401k plan
  • Discretionary Paid Time Off and Paid Holidays
  • Parental Leave 
  • Equity 
  • Monthly Wellness Reimbursement
  • Monthly Lunch on Legion

ABOUT LEGION

Join Legion's mission to turn hourly jobs into good jobs. We're a remote, mission-driven team seeking exceptional talent to propel this vision. Embrace a culture that's collaborative, fast-paced, and entrepreneurial. With us, you'll grow your skills, work closely with experienced executives, and contribute significantly to our mission.

Legion Technologies delivers the industry’s most innovative workforce management platform. It enables businesses to maximize labor efficiency and employee engagement simultaneously. The award-winning, AI-driven Legion WFM platform is intelligent, automated, and employee-centric. It’s proven to deliver 13x ROI through schedule optimization, reduced attrition, increased productivity, and increased operational efficiency. Legion delivers cutting-edge technology in an easy-to-use platform and mobile app that employees love. 

If you're ready to make an impact and grow your career, Legion is where you belong. Join us in making hourly work rewarding and fulfilling.

BACKGROUND AND OPPORTUNITY 

There are almost 75 million hourly workers in the United States, representing more than half of the entire workforce. Historically, managing hourly employees has been difficult due to high attrition (average of 60%) and high replacement costs (average of $3,200 per employee in retail). The ongoing labor shortage and competition from the gig economy make it more difficult to attract and retain hourly employees. The top reasons hourly employees leave their jobs are a lack of schedule empowerment, poor communication with employers, and an inability to get paid early. Gen Z and the millennial workforce demand gig-like flexibility, modern technology, and compelling work options. Legion’s mission is to turn hourly jobs into good jobs, serving the hourly workers who make up the majority of the US workforce. We believe in empowering employees and helping employers be efficient and innovative by enabling intelligent automation powered by Legion’s Workforce Management platform to optimize labor efficiency and enhance the employee experience simultaneously. Legion WFM was built for the cloud with AI at the core and designed to handle the complexity of modern businesses and meet the needs of today’s hourly employees.  Our team is comprised of dedicated individuals from all backgrounds and experiences, globally distributed across all time zones.

For more information, visit https://legion.co 

EQUAL EMPLOYMENT OPPORTUNITY

Legion Technologies is proud to be an equal-opportunity employer and is committed to maintaining a diverse and inclusive work environment. All qualified applicants will be considered for employment without regard to race, color, religion, sex, age, disability, marital status, familial status, sexual orientation, pregnancy, genetic information, gender identity, gender expression, national origin, ancestry, citizenship status, veteran status, and any other legally protected status under federal, state, or local anti-discrimination laws.

DISABILITY ACCOMMODATION

For individuals with disabilities who need additional assistance at any point in the application and interview process, please email recruiting@legion.co 

We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Legion recruiting email communications will always come from the @legion.co domain. Any outreach claiming to be from Legion via other sources should be ignored. If you are uncertain whether you have been contacted by an official Legion employee, reach out to recruiting@legion.co

 

Legion is an equal opportunity employer. All applicants will be considered for employment without attention to race, religion, color, sex, sexual orientation, gender identity, age, national origin, veteran, disability status, or any other basis covered by appropriate law.

How We Determine What We Pay

As a global employer, Legion determines pay for positions using local, national, and industry-specific survey data. We evaluate external equity and the cost of labor/prevailing wage index in the relative marketplace for jobs directly comparable to jobs within our company. Our posted salary range is based on national data and may be refined for a candidate's region/town/cost of living. For new hires, we strive to make competitive offers allowing the new employee room for future growth. Salaries will be based on the applicant’s location, level of experience, education, and specialized knowledge and skills. Additionally, we consider the external market rate, the amount we have budgeted internally, and internal equity within the company for the same position. An employee/candidate with a stronger skill set will receive higher pay.

Job Applicant Privacy Policy

This Job Applicant Privacy Policy (“Policy”) describes how Legion Technologies, Inc. (“Legion”, “we”, “us” and “our”) collects, uses, and discloses “personal information” as defined under California law from and about job applicants who are residents of California.

This Policy does not apply to our handling of data gathered about you in your role as a user of our consumer-facing services. When you interact with us as in that role, the Legion Privacy Policy applies.

  1. Types of Personal Information We Handle

    We collect, store, and use various types of personal information through the application and recruitment process. We collect such information either directly from you or (where applicable) from another person or entity, such as an employment agency or consultancy, background check provider, or other referral sources. This information includes:

    • Identification and contact information, and related identifiers such as full name, date and place of birth, citizenship and permanent residence, home and business addresses, telephone numbers, email addresses, and such information about your beneficiaries or emergency contacts.
    • Professional or employment-related information, including:
      • Recruitment, employment, or engagement information such as application forms and information included in a resume, cover letter, or otherwise provided through any application or engagement process; and copies of identification documents, such as driver’s licenses, passports, and visas; and background screening results and references.
      • Career information such as job titles; work history; work dates and work locations; information about skills, qualifications, experience, publications, speaking engagements, and preferences; and professional memberships
    • Education Information such as institutions attended, degrees, certifications, training courses, publications, and transcript information.
    • Legally protected classification information such as race, sex/gender, religious/ philosophical beliefs, gender identity/expression, sexual orientation, marital status, military service, nationality, ethnicity, request for family care leave, political opinions, and criminal history.
    • Other information such as any information you voluntarily choose to provide in connection with your job application.
  2. How We Use Personal Information

    We collect, use, share, and store personal information from job applicants for our and our service providers’ business and operational purposes in the recruitment process such as: processing your application, tracking your application through the recruitment process, contacting references with your authorization, conducting background checks you authorize, and making hiring decisions. We will also use job applicant information for internal analysis purposes to understand the applicants who apply and to improve our recruitment process. We may sometimes need to use applicant information for legal purposes, such as in connection with any challenges made to our hiring decisions.

  3. With Whom We Share Personal Information

    We will disclose job applicant personal information to the following types of entities or in the following circumstances (where applicable):

    • Internally: to other Legion personnel involved in the recruiting and hiring process.
    • Vendors: such as technology service providers, travel management providers, human resources suppliers, background check companies, and employment agencies or recruiters, where applicable.
    • Legal Compliance: when required to do so by law, regulation, or court order or in response to a request for assistance by the police or other law enforcement agency.
    • Litigation Purposes: to seek legal advice from our external lawyers or in connection with litigation with a third party.
    • Business Transaction Purposes: in connection with the sale, purchase, or merger.
  4. How to Contact Us About this Policy – If you have any questions about this Policy, please contact privacy@legion.co.

To apply: https://weworkremotely.com/remote-jobs/legion-director-of-production-engineering

Read the full description
Engineer Principal Software Engineer at HubSpot

Designs and builds HubSpot's observability platform for distributed systems and AI agents, setting architecture patterns and telemetry standards across hundreds of microservices.

Lead Posted about 16 hours ago RemoteFirstJobs Product
What this role involves

POS-5690

About the Team

The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices, billions of daily events, and thousands of engineers who depend on that signal to ship reliably.

We are now investing in the next generation of this platform. As HubSpot deploys AI agents and ML-powered features across the product, the team is building the tracing and telemetry primitives that make it possible to understand, debug, and trust what those systems are doing in production. This is greenfield, technically interesting work at a scale few companies operate at and we are looking for a Principal Engineer to help lead it.

About the Role

We are seeking a Principal Software Engineer to be the technical anchor for HubSpot’s Observability platform. This role sits at the intersection of large-scale distributed systems, developer platform design, and AI observability. A big part of this role is working horizontally across a large engineering org and setting patterns and standards that make it easier for teams to instrument, alert on, and reason about their services. You will also shape how we trace and understand our growing fleet of AI agents and ML systems in production: a technically distinct and increasingly critical problem.

Key Expectations

  • Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization.
  • AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production. Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work.
  • High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency. Make principled tradeoffs between sampling, fidelity, retention, and developer ergonomics.
  • Hands-on, High-Leverage Builder: Ship production code. Lead design reviews and take high-impact initiatives end-to-end, from prototype to production system at scale. Stay close to the systems you build and be the person who can debug the hardest problems when they surface.
  • Developer Experience & Adoption: Design the instrumentation APIs and libraries that product engineers reach for, making correct observability the path of least resistance. Drive OpenTelemetry adoption across a large, polyglot codebase. Build the tooling that turns raw telemetry into actionable signal for teams operating at speed.
  • Production Intelligence & Reliability Patterns: Define patterns for SLO/SLI design, alerting philosophy, and how teams graduate from reactive to proactive incident response. Push for simplicity in a domain that wants to get complicated, and consistency where tooling can drift across a large organization.
  • Technical Leadership & Influence: Partner with infrastructure, platform, and product engineering teams to understand their signal gaps and close them. Influence technical strategy alongside engineering leadership, translating observability constraints and opportunities into product and operational decisions. Mentor senior engineers and tech leads, driving thoughtful design decisions and capturing learnings from major incidents and large-scale migrations.

What You Bring

  • Platform-Builder Experience: Proven experience building observability or telemetry tooling for internal engineering teams, rather than simply consuming it. You understand how to architect developer platforms that serve thousands of engineers across a large organization, backed by deep operational instincts and hard-earned expertise.
  • Telemetry Systems Depth: Deep expertise navigating trade-offs in telemetry pipeline design across high-cardinality data, dynamic sampling, query latency, retention economics, and data fidelity. Strong technical fluency with OpenTelemetry, distributed tracing engines, metric backends, and large-scale log ingestion infrastructure.
  • Incident Automation & Operational Excellence: Proven track record linking telemetry signals directly to automated operational workflows. You have designed architecture for real-time telemetry triggers that power automated remediation, dynamic runbooks, or AI-assisted root-cause diagnosis across microservices environments.

Why This Role, Why Now

HubSpot is scaling fast — more engineers, more microservices, more AI systems running in production — and the Observability platform is at an inflection point. The foundations are solid, but the next phase requires a different kind of investment: rethinking cardinality economics, making OpenTelemetry the default across a large polyglot org, and building an entirely new layer of AI tracing that doesn’t yet exist.

This Principal Engineer will have a direct line of sight from the architecture they design to the engineering outcomes we measure: incident response times, adoption rates, developer satisfaction, and the trust that product teams place in their production signal. The AI observability layer in particular is greenfield — there is no playbook to follow, which is exactly what makes this the right moment for the right person.

If you want to build the platform that helps thousands of engineers understand what their systems are doing — including systems powered by AI that are genuinely hard to see inside — this is the role.

Pay & Benefits

The cash compensation below includes base salary, on-target commission for employees in eligible roles, and annual bonus targets under HubSpot’s bonus plan for eligible roles. In addition to cash compensation, some roles are eligible to participate in HubSpot’s equity plan to receive restricted stock units (RSUs). Some roles may also be eligible for overtime pay. Individual compensation packages are tailored to your skills, experience, qualifications, and other job-related reasons.

This resource will help guide how we recommend thinking about the range you see. Learn more about HubSpot’s compensation philosophy.

Benefits are also an important piece of your total compensation package. Explore the benefits and perks HubSpot offers to help employees grow better.

At HubSpot, fair compensation practices aren’t just about checking off the box for legal compliance. It’s about living out our value of transparency with our employees, candidates, and community.

Annual Cash Compensation Range:

$313,800—$502,080 USD

At HubSpot, we value both flexibility and connection. Whether you’re a Remote employee or work from the Office, we want you to start your journey here by building strong connections with your team and peers. If you are joining our Engineering team, you will be required to attend a regional HubSpot office for in-person onboarding. If you join our broader Product team, you’ll also attend other in-person events, such as your Product Group Summit and other gatherings, to continue building on those connections.

If you require an accommodation due to travel limitations or other reasons, please inform your recruiter during the hiring process. We are committed to supporting candidates who may need alternative arrangements

About HubSpot

HubSpot (NYSE: HUBS) is an AI-powered customer platform with all the software, integrations, and resources customers need to connect marketing, sales, and service. HubSpot’s connected platform enables businesses to grow faster by focusing on what matters most: customers.

At HubSpot, bold is our baseline. Our employees around the globe move fast, stay customer-obsessed, and win together. Our culture is grounded in four commitments: Solve for the Customer, Be Bold, Learn Fast, Align, Adapt & Go!, and Deliver with HEART. These commitments shape how we work, lead, and grow.

We’re building a company where people can do their best work. We focus on brilliant work, not badge swipes. By combining clarity, ownership, and trust, we create space for big thinking and meaningful progress. And we know that when our employees grow, our customers do too.

Recognized globally for our award-winning culture by Comparably, Glassdoor, Fortune, and more, HubSpot is headquartered in Cambridge, MA, with employees and offices around the world.

Explore more:

  • HubSpot Careers
  • Life at HubSpot on Instagram

If you need accommodations or assistance due to a disability, please reach out to us using this form.

Massachusetts Applicants: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Germany Applicants: (m/f/d) - link to HubSpot’s Career Diversity page here.

India Applicants: link to HubSpot India’s equal opportunity policy here.

HubSpot may use AI to help screen or assess candidates, but all hiring decisions are always human. More information can be found here. By submitting your application, you agree that HubSpot may collect your personal data for recruiting, global organization planning, and related purposes. We may use CLEAR ID Verification during the hiring process to confirm your identity and help maintain a safe, secure, and trusted experience for all candidates. Refer to HubSpot’s Recruiting Privacy Notice for details on data processing and your rights.

Read the full description
Engineer Staff Site Reliability Engineer at Replit

Staff SRE architecting observability solutions, defining reliability standards, leading incident response, and automating infrastructure at scale.

Lead Posted about 16 hours ago RemoteFirstJobs Product
What this role involves

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the role:

Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit’s infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.

We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure’s reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit.

You Will:

  • Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection.

  • Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed.

  • Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR).

  • Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios.

  • Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions.

  • Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose.

  • Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs.

  • Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture.

  • Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it’s building new internal tools or integrating with third-party vendors.

Required Skills and Experience:

  • 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).

  • Strong programming skills in languages like Python or Go. You write high-quality, well-tested code.

  • Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture.

  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.

  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing).

  • Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure.

  • Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.

  • Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices.

  • Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels.

  • A willingness to dive into understanding, debugging, and improving any layer of the stack.

  • You’re passionate about making software creation accessible and empowering the next generation of builders.

Bonus Points:

  • Deep experience with Google Cloud Platform (GCP) services and tools.

  • Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry).

  • Experience designing and building reliable systems capable of handling high throughput and low latency.

  • Significant experience with Go and Terraform.

  • Familiarity with working in rapid-growth, startup environments.

  • Experience writing company-facing blog posts and training materials.

Full-Time Employee Benefits Include:

💰 Competitive Salary & Equity

đŸ’č 401(k) Program with a 4% match ( US Only)

⚕ Health, Dental, Vision and Life Insurance

đŸ©Œ Short Term and Long Term Disability

đŸšŒ Paid Parental, Medical, Caregiver Leave

🏝 Flexible Time Off (FTO) + Holidays

🚗 Commuter Benefits ( In-Office & US Only)

đŸ“± Monthly Wellness Stipend

đŸ§‘â€đŸ’» Autonomous Work Environment

đŸ–„ In Office Set-Up Reimbursement ( In-Office Only)

🚀 Quarterly Team Gatherings

☕ In Office Amenities ( In-Office Only)

Want to learn more about what we are up to?

  • Self-driving Company

  • Replit Agent at Scale

  • AI Adoption

  • Build Open-Source Apps

Interviewing + Culture at Replit

  • Operating Principles

  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Read the full description
Engineer Principal Software Engineer at HubSpot

Designs and leads HubSpot's observability platform architecture, including distributed tracing, metrics, logging, and AI/ML system telemetry at scale across hundreds of microservices.

Lead Posted about 16 hours ago RemoteFirstJobs Product
What this role involves

POS-5690

About the Team

The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices, billions of daily events, and thousands of engineers who depend on that signal to ship reliably.

We are now investing in the next generation of this platform. As HubSpot deploys AI agents and ML-powered features across the product, the team is building the tracing and telemetry primitives that make it possible to understand, debug, and trust what those systems are doing in production. This is greenfield, technically interesting work at a scale few companies operate at and we are looking for a Principal Engineer to help lead it.

About the Role

We are seeking a Principal Software Engineer to be the technical anchor for HubSpot’s Observability platform. This role sits at the intersection of large-scale distributed systems, developer platform design, and AI observability. A big part of this role is working horizontally across a large engineering org and setting patterns and standards that make it easier for teams to instrument, alert on, and reason about their services. You will also shape how we trace and understand our growing fleet of AI agents and ML systems in production: a technically distinct and increasingly critical problem.

Key Expectations

  • Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization.
  • AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production. Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work.
  • High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency. Make principled tradeoffs between sampling, fidelity, retention, and developer ergonomics.
  • Hands-on, High-Leverage Builder: Ship production code. Lead design reviews and take high-impact initiatives end-to-end, from prototype to production system at scale. Stay close to the systems you build and be the person who can debug the hardest problems when they surface.
  • Developer Experience & Adoption: Design the instrumentation APIs and libraries that product engineers reach for, making correct observability the path of least resistance. Drive OpenTelemetry adoption across a large, polyglot codebase. Build the tooling that turns raw telemetry into actionable signal for teams operating at speed.
  • Production Intelligence & Reliability Patterns: Define patterns for SLO/SLI design, alerting philosophy, and how teams graduate from reactive to proactive incident response. Push for simplicity in a domain that wants to get complicated, and consistency where tooling can drift across a large organization.
  • Technical Leadership & Influence: Partner with infrastructure, platform, and product engineering teams to understand their signal gaps and close them. Influence technical strategy alongside engineering leadership, translating observability constraints and opportunities into product and operational decisions. Mentor senior engineers and tech leads, driving thoughtful design decisions and capturing learnings from major incidents and large-scale migrations.

What You Bring

  • Platform-Builder Experience: Proven experience building observability or telemetry tooling for internal engineering teams, rather than simply consuming it. You understand how to architect developer platforms that serve thousands of engineers across a large organization, backed by deep operational instincts and hard-earned expertise.
  • Telemetry Systems Depth: Deep expertise navigating trade-offs in telemetry pipeline design across high-cardinality data, dynamic sampling, query latency, retention economics, and data fidelity. Strong technical fluency with OpenTelemetry, distributed tracing engines, metric backends, and large-scale log ingestion infrastructure.
  • Incident Automation & Operational Excellence: Proven track record linking telemetry signals directly to automated operational workflows. You have designed architecture for real-time telemetry triggers that power automated remediation, dynamic runbooks, or AI-assisted root-cause diagnosis across microservices environments.

Why This Role, Why Now

HubSpot is scaling fast — more engineers, more microservices, more AI systems running in production — and the Observability platform is at an inflection point. The foundations are solid, but the next phase requires a different kind of investment: rethinking cardinality economics, making OpenTelemetry the default across a large polyglot org, and building an entirely new layer of AI tracing that doesn’t yet exist.

This Principal Engineer will have a direct line of sight from the architecture they design to the engineering outcomes we measure: incident response times, adoption rates, developer satisfaction, and the trust that product teams place in their production signal. The AI observability layer in particular is greenfield — there is no playbook to follow, which is exactly what makes this the right moment for the right person.

If you want to build the platform that helps thousands of engineers understand what their systems are doing — including systems powered by AI that are genuinely hard to see inside — this is the role.

Pay & Benefits

The cash compensation below includes base salary, on-target commission for employees in eligible roles, and annual bonus targets under HubSpot’s bonus plan for eligible roles. In addition to cash compensation, some roles are eligible to participate in HubSpot’s equity plan to receive restricted stock units (RSUs). Some roles may also be eligible for overtime pay. Individual compensation packages are tailored to your skills, experience, qualifications, and other job-related reasons.

This resource will help guide how we recommend thinking about the range you see. Learn more about HubSpot’s compensation philosophy.

Benefits are also an important piece of your total compensation package. Explore the benefits and perks HubSpot offers to help employees grow better.

At HubSpot, fair compensation practices aren’t just about checking off the box for legal compliance. It’s about living out our value of transparency with our employees, candidates, and community.

Annual Cash Compensation Range:

$313,800—$502,080 USD

At HubSpot, we value both flexibility and connection. Whether you’re a Remote employee or work from the Office, we want you to start your journey here by building strong connections with your team and peers. If you are joining our Engineering team, you will be required to attend a regional HubSpot office for in-person onboarding. If you join our broader Product team, you’ll also attend other in-person events, such as your Product Group Summit and other gatherings, to continue building on those connections.

If you require an accommodation due to travel limitations or other reasons, please inform your recruiter during the hiring process. We are committed to supporting candidates who may need alternative arrangements

About HubSpot

HubSpot (NYSE: HUBS) is an AI-powered customer platform with all the software, integrations, and resources customers need to connect marketing, sales, and service. HubSpot’s connected platform enables businesses to grow faster by focusing on what matters most: customers.

At HubSpot, bold is our baseline. Our employees around the globe move fast, stay customer-obsessed, and win together. Our culture is grounded in four commitments: Solve for the Customer, Be Bold, Learn Fast, Align, Adapt & Go!, and Deliver with HEART. These commitments shape how we work, lead, and grow.

We’re building a company where people can do their best work. We focus on brilliant work, not badge swipes. By combining clarity, ownership, and trust, we create space for big thinking and meaningful progress. And we know that when our employees grow, our customers do too.

Recognized globally for our award-winning culture by Comparably, Glassdoor, Fortune, and more, HubSpot is headquartered in Cambridge, MA, with employees and offices around the world.

Explore more:

  • HubSpot Careers
  • Life at HubSpot on Instagram

If you need accommodations or assistance due to a disability, please reach out to us using this form.

Massachusetts Applicants: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Germany Applicants: (m/f/d) - link to HubSpot’s Career Diversity page here.

India Applicants: link to HubSpot India’s equal opportunity policy here.

HubSpot may use AI to help screen or assess candidates, but all hiring decisions are always human. More information can be found here. By submitting your application, you agree that HubSpot may collect your personal data for recruiting, global organization planning, and related purposes. We may use CLEAR ID Verification during the hiring process to confirm your identity and help maintain a safe, secure, and trusted experience for all candidates. Refer to HubSpot’s Recruiting Privacy Notice for details on data processing and your rights.

Read the full description
Engineer Senior Staff Software Engineer, Partner Integrations at Flex

Senior Staff Software Engineer oversees technical roadmap for partner integrations, building APIs and SDKs that enable external partners to connect to Flex's rent payment platform.

Lead Posted about 16 hours ago RemoteFirstJobs Product
What this role involves

Flex is a growth-stage, NYC headquartered FinTech company that is creating the best rent payment experience. It’s hard to believe that it’s 2026 and paying rent on time is expensive, inflexible, and difficult. We’re here to change that! Flex enables our users to pay rent throughout the month on a schedule that better fits their finances and budget. Our mission is to empower as many renters as possible with flexibility over their most significant recurring expense. After deliberately keeping a stealth profile as we built up unprecedented investor support and an enthusiastic user base, we are looking for motivated individuals to help us keep our mission growing. Will you be a part of the team?

About Our Opportunity

Flex exists to make paying rent work the way real life actually works — smoothing out cash flow, helping renters avoid late fees, and letting them build credit instead of falling behind. As a Senior Staff Software Engineer, you’ll oversee the technical roadmap for the team, working across APIs, SDKs, and Web experiences. You will work with teams across the organization including Engineering, Product, Design, Infrastructure, Sales, Partner and Customer Success to ensure that the technical strategy meets our goals. We expect you to be hands-on and execute work as an individual, and build products that allow for flexibility as we evolve our product offerings.

About Our Team!

We have some really exciting teams who are looking for an amazing engineer like you! Check them out!

  • Partner Integrations - Partner integrations work to standardize APIs, tools, and processes that let external partners connect to Flex for data sharing and payments across entry points like Flex Anywhere

Who thrives here?

  • A Doer— you get things done efficiently, and you’re comfortable rolling up your sleeves when something’s blocking the team.
  • An Owner — you take full accountability for what you build, from design through production, and you always have a plan B.
  • Someone Collaborative— you bring people together to pressure-test ideas rather than working in a silo.
  • Someone with Precision — your code and your communication are both clear, because ideas only matter once others can act on them.
  • Someone Resilient— systems break and priorities shift, and you adapt without losing momentum.
  • Someone Humble— you care more about getting to the right answer than being the one who’s right.

Qualifications:

  • Bachelor or above degree in Computer Science or a related field
  • Minimum of 10 years experience in software engineering, with at least 3 years of technical leadership experience in a hands-on capacity
  • Ability to work on a globally-distributed team with a high degree of ownership
  • Experience leading the delivery of multiple highly impactful products end to end, on time with a high quality bar
  • Experience working with technical and non-technical stakeholders, successfully aligning and setting expectations on scope and delivery
  • Ability to drive yourself and your team to bring quality and consistency to their code and architecture, without compromising velocity
  • Ability to grow in a fast-paced and dynamic environment that will challenge you to always bring your best
  • Experience in designing and developing solutions that maximize ROI and lead to significant business impact
  • Experience working in FinTech and familiar with major payment rails
  • Experience working with Sales stakeholders and/or building products for Sales professionals
  • Experience building integrations with external partners & managing relationships with external stakeholders
  • Experience working on AWS cloud based applications

Compensation

Flex takes a market-based approach to pay, and compensation may vary depending on your primary work location. Work locations are categorized into one of three tiers based on a cost of labor index for that geographic area. The successful candidate’s starting pay will be commensurate with their experience, qualifications, and Flex’s internal leveling guidelines and benchmarks.

Tier 1 (NYC/Bay Area, Los Angeles, Seattle)

$240,000—$300,000 USD

Tier 2 (Austin, Washington D.C. Philadelphia, San Diego, Chicago, Atlanta)

$216,000—$270,000 USD

Tier 3 (Salt Lake City, all other USA cities)

$204,000—$255,000 USD

Life at Flex

We understand that it takes a diverse team of highly intelligent, curious, determined, empathetic, and self aware people to grow a successful company. Our HQ is located in New York City, but we have employees located throughout the US, Australia, Canada and South America. We are growing quickly, but deliberately, with a focus on building an inclusive culture. Our dynamic team has incredible perspectives to share, just as we know you do, and we take great pride in being an equal opportunity workplace.

Offices

Roles posted in New York, San Francisco, and Salt Lake City are hybrid positions with on-site expectations of 2-3 days per week in our local offices. For candidates outside of these areas, you may be eligible for our relocation assistance program.

Benefits

For full-time U.S. employees we offer:

  • Competitive medical, dental, and vision
  • Company equity
  • 401(k) plan with company match
  • Unlimited paid time off + 13 company paid holidays
  • Parental leave
  • Free Flex subscription

For full-time non-U.S. employees, we offer:

  • Competitive compensation + company equity
  • Unlimited PTO
Read the full description
Engineer Staff Software Backend Engineer, Tech Lead – Foundations

Technically leads the Foundations platform team, building shared backend infrastructure and primitives that other engineering teams depend on.

Lead Posted about 16 hours ago Jobicy AI
What this role involves
NetBox Labs seeks a Staff Software Backend Engineer to technically lead our Foundations team. Foundations is the platform team that builds shared platform primitives that other engineering teams depend on...
Read the full description
Engineer Director – Forward Deployed Engineering at Re:Build Manufacturing

Leads a team of forward-deployed engineers building full-stack digital solutions across manufacturing sites, managing technical strategy, delivery, and people development.

Lead Posted 1 day ago RemoteFirstJobs Product
What this role involves

About Re:Build

At Re:Build, our mission is to ensure the next generation of important products are made, at scale, in America. We are laying the foundation for a better future for our customers, employees, and communities by revitalizing America’s manufacturing base and creating meaningful jobs across the country, including in historically deindustrialized regions.

We operate an advanced, end-to-end manufacturing platform that partners with industrial companies and innovators to take products from first concept to full-scale production in critical verticals including aerospace and defense, electrification, medical, energy and environment, and robotics and automation.

The way we operate is as important as the work we do. It’s guided by The Re:Build Way, 16 principles that shape how we collaborate with each other, partner with our customers and vendors, and contribute to the communities where we operate. (link to The Re:Build Way principles )

Who we are looking for

The Director, Forward Deployed Engineering is a hands-on leadership role in Re:Build Manufacturing’s Digital Innovation Group (DIG). This role reflects the organization’s belief that digital innovation is key to scaling as a premier industrial company. In this position, the Director is responsible for hiring, leading, and developing a team of Forward Deployed Engineers (FDEs) who are embedded across Re:Build sites and Resource Center functions, implementing full-stack digital solutions that solve real-world business problems and drive measurable business impact. Success in the role depends on the ability to build trust, ensure alignment, and partner effectively with internal stakeholders throughout the entire product lifecycle. The ideal candidate combines a strong software engineering background with proven experience managing software delivery and solution engineering that directly serves customers. They bring deep technical credibility, strong people leadership skills, and the judgment needed to keep engagements well prioritized and delivered quickly.

What you get to do

  • Partner with portfolio Manufacturing, Engineering Services, and Functional teams to identify difficulties, operational inefficiencies, and opportunities for digital solutions
  • Evaluate and prioritize digital solution opportunities based on business impact, feasibility, and strategic alignment
  • Guide FDEs through sophisticated technical decisions, feature development, fixing issues, code reviews, and delivery execution across collaborator engagement
  • Define and implement engineering standards and coding practices across FDE projects to ensure scalability, maintainability, and security of delivered solutions
  • Build the repeatable delivery machine of reusable assets, playbooks, user documentation, and implementation approaches that strengthen delivery excellence and collaborator enablement
  • Establish metrics to measure solution impact and value post-deployment and gather feedback for continuous improvement
  • Evaluate and foster the adoption of emerging AI-assisted development tools (e.g., code generation, review, testing automation) across the FDE team to accelerate delivery velocity and code quality
  • Know the latest digital trends in AI, software development, manufacturing, Industry 4.0 technologies, and relevant SaaS solutions
  • Ability to travel 25% to customer sites (will only require domestic US travel)

What you bring to the Team

  • Bachelor’s degree in Engineering, Computer Science, or related field required; MBA or relevant advanced degree; or equivalent combination of education and experience preferred
  • 8+ years of experience leading customer-facing software delivery, implementation, solution engineering, consulting, or technical delivery programs within enterprise SaaS, technology, or systems integrator environments
  • 2+ years of people leadership or technical leadership experience leading and developing engineers, architects, consultants, or technical delivery teams
  • 3-5 years of hands-on software development, solution engineering, integration development, or full-stack development experience, including technical leadership responsibilities
  • Experience with DevSecOps practices, secure software development lifecycle (SDLC) methodologies, and Agile/SAFe delivery frameworks
  • Fluency in AI technologies applied to software development lifecycle
  • Proficiency with project and delivery management tools such as Jira and Confluence, with a proven ability to lead cross-functional teams and multiple concurrent engagements
  • Experience in manufacturing operations, industrial engineering, or engineering services strongly preferred

The BIG payoff

We are a company who is going to make a difference in the industries and the communities in which we choose to operate.  Every employee of Re:Build will share ownership in the company and will share in the financial rewards of the success we achieve together, at all levels of the company!

We want to work with people that reflect the communities in which we operate

Re:Build Manufacturing is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, marital status, parental status, cultural background, organizational level, work styles, tenure and life experiences. Or for any other reason.

Re:Build is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or an accommodation due to a disability, you may contact us at accommodations.ta@ReBuildmanufacturing.com or you may call us at 617.909.6275.

Read the full description
Engineer Lead Data Engineer at Box

Designs and manages enterprise data platforms, lakehouse architectures, and data pipelines while leading teams on data governance, quality, and compliance.

Lead Posted 1 day ago RemoteFirstJobs Product
What this role involves

\*\*\* This is where your organization can create a consistent intro to all of your jobs, creating consistency in voice and messaging across all job posts

\*\*\* C’est ici que votre organisation peut crĂ©er une introduction cohĂ©rente Ă  tous vos emplois, en crĂ©ant une cohĂ©rence dans la voix et la messagerie dans tous les postes.

Marlabs, a global AI and Digital Solutions Consulting firm, delivers intelligent solutions across AI, data, analytics, and product engineering. Since 2000, we have partnered with some of the largest healthcare, life sciences, financial services, and government organizations worldwide. As we continue to expand our global footprint, we have an exciting opportunity for a highly skilled Lead Data Engineer to join our innovative and dynamic team.

Lead Data Engineer | About You

As a Lead Data Engineer, you will be responsible for designing, building, and managing the organization’s modern data platform, ensuring reliable, secure, and scalable data products that support analytics, reporting, and AI-driven business initiatives. You will lead the development of enterprise data pipelines, lakehouse architecture, and governance frameworks while partnering closely with AI/ML, platform engineering, and security teams. The ideal candidate combines deep expertise in data engineering, data modeling, cloud-based architectures, and data governance with a strong focus on reliability, observability, and regulatory compliance.

Lead Data Engineer | Day-to-Day

  • Design, build, and maintain scalable lakehouse architectures and enterprise data platforms that support analytics, reporting, and AI-driven solutions.
  • Develop and manage secure data ingestion frameworks, including CDC, batch, API, and file-based integrations from operational and transactional source systems.
  • Create and maintain data models, semantic layers, and governed metrics that enable consistent, trusted, and business-ready data consumption.
  • Implement data quality, reconciliation, observability, and monitoring processes to ensure reliable, recoverable, and high-performing data pipelines.
  • Partner with AI/ML, platform engineering, and security teams to deliver governed data products, support regulatory compliance requirements, and ensure proper data classification and access controls.
  • Establish and enforce data governance standards, lineage documentation, data contracts, quality thresholds, and operational procedures while supporting onboarding of new data sources and environments.

Lead Data Engineer | Skills & Experience

  • 8+ years of experience in Data Engineering, including end-to-end ownership of data ingestion, transformation, storage, and analytics delivery, with experience leading large-scale data initiatives.
  • Strong expertise in Python and SQL, including advanced data modeling, transformation frameworks, data quality management, and performance optimization.
  • Hands-on experience with modern lakehouse architectures and data platforms, including technologies such as Apache Iceberg, Delta Lake, Hudi, object storage, and cloud-native data solutions.
  • Proven experience building scalable data pipelines and CDC solutions, leveraging technologies such as Kafka, Debezium, Airflow, Dagster, dbt, and enterprise integration frameworks.
  • Strong understanding of data governance, lineage, security, and compliance practices, including data contracts, access controls, observability, audit readiness, and regulated industry environments.
  • Experience collaborating with AI/ML, analytics, platform engineering, and business teams to deliver trusted, governed, and scalable data products; financial services or banking industry experience is highly preferred.

\*\*\* Similar to the introduction that can precede all job descriptions, an outro can also be formatted for consistency on all posts

\*\*\* Semblable Ă  l’introduction qui peut prĂ©cĂ©der toutes les descriptions de poste, une outro peut Ă©galement ĂȘtre formatĂ©e pour la cohĂ©rence sur tous les messages

Read the full description
Engineer Staff Software Engineer at Fluxon

Staff-level engineer who writes production code, leads technical projects, mentors engineers, and shapes architectural decisions across multiple technology stacks.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Who we are

At Fluxon, we believe that how you build matters as much as what you build. We help businesses navigate their most important technology decisions with confidence, and take responsibility for seeing them through. Founded by ex-Googlers and startup veterans, we’re proud to partner with teams behind some of the most ambitious products, including Google, OpenAI, Anthropic, Walmart and Stripe.

Our work spans strategy, design, and engineering — often in complex, AI-driven environments — where clarity, speed and quality are the standard. We use AI intentionally, applying it only where it adds real value and expands what’s possible. Care shapes everything we do.

Inside Fluxon, you’ll find a global, remote-first team of experienced builders, who are curious, kind and serious about their craft. We’re building a place where people can take ownership, solve problems that matter and do work they’re proud to stand behind. If you want to do your best work alongside people who care as much as you do, you’ll feel at home here.

This role is fully remote, with candidates based in Buenos Aires, Argentina.

About the role

As a Staff Software Engineer at Fluxon, you’ll play a key role in shaping the technical direction of our engineering organization. This is a highly senior, hands-on leadership position where you’ll partner closely with company and engineering leadership to influence strategy, guide architectural decisions, and elevate our overall engineering practice. All Staff Engineers write production code, and everyone joins Fluxon as an individual contributor before stepping into project leadership or management.

You’ll be responsible for:

  • Guiding product delivery all the way to the user, leading projects, providing technical guidance, and building and iterating in a dynamic environment
  • Partnering directly with clients to understand their needs and achieve business goals
  • Defining product requirements, identifying appropriate system designs and planning development in partnership with our Product and Design teams
  • Helping drive a healthy and effective engineering culture within customer teams and inside Fluxon
  • Mentoring engineers across multiple teams, supporting their ongoing growth and strengthening team capabilities

You’ll work with a diversity of technologies, including:

  • Core Languages & Runtimes

    • Primary Languages: TypeScript/JavaScript, Python, Golang (Go), Java, C# (.NET), Kotlin, Swift, Rust
    • Additional: Ruby on Rails, Java, C# (.NET), Kotlin, Swift, Rust
  • Frameworks & Ecosystems

    • Front-End & UI: React, Next.js (Full-Stack), Angular, SwiftUI
    • Back-End & Server-Side: Spring Boot (Java/Kotlin), FastAPI (Python), Django (Python)
    • Mobile Development: Expo (React Native)
  • Cloud & Infrastructure

    • Platforms: Google Cloud Platform (GCP), Amazon Web Services (AWS), Microsoft Azure
    • Compute: GCP Compute Engine (VMs), AWS Fargate (Container Orchestration), Google Cloud Run (Serverless Containers), AWS Amplify
    • Storage: AWS S3 (Object Storage), Google Cloud Storage (GCS)
  • Data & Messaging Services

    • Streaming & Queuing: Apache Kafka (Event Streaming), AWS SQS (Simple Queue Service)
    • Data Warehouse & Analytics: Google BigQuery
    • Monitoring & Observability: GCP Cloud Monitoring Suite (CMS)
  • Data Stores

    • Relational (SQL): PostgreSQL, MariaDB
    • NoSQL/Document: Firestore (Firebase), Supabase, MongoDB
    • In-Memory Caching: Redis, Memcache
  • Advanced Technologies & Architecture

    • Artificial Intelligence/Machine Learning: AI/ML, Large Language Models (LLMs), Agentic AI, Natural Language Processing (NLP)
    • LLM Platforms: Google Gemini, OpenAI ChatGPT, Vertex AI (GCP), Anthropic Claude, Hugging Face (OSS Models)
    • Software Design: Architecture Redesign, Single Page Applications (SPA), Mobile Application Development
    • Emerging Tech: Blockchain/Crypto

Qualifications

  • 7+ years of industry experience in software development
  • Experience leading development through the full product lifecycle, including CI/CD, testing, release management, deployment, monitoring and incident response
  • Fluent in the design and implementation of scalable system architectures, data structures and algorithms, and effective development practices
  • Professional fluency in written and spoken English, with the ability to communicate technical concepts clearly to both international teammates and customers

What we offer

  • Remote-first, flexible work with a budget to set up a work space that works for you
  • Localized healthcare coverage to support you and your family’s wellbeing
  • Flexible paid time off with a minimum of 3 weeks per year (plus holidays)
  • “No internal meetings” Fridays, so you can focus on deep, uninterrupted work
  • A professional growth budget for learning that matters to you –  whether it’s developing your technical skills or learning a new language
  • A monthly wellness allowance to support your physical and mental health
  • Annual company offsites, where we gather in person to build meaningful connections
  • A competitive salary that reflects your expertise, impact and experience
  • Profit-sharing, so you can benefit directly from the value we build together
  • A paid sabbatical program, designed for rest, renewal and fresh perspective

We believe diverse teams perform better, and an inclusive environment is essential to building a successful organization. We welcome applicants from all backgrounds, experiences, and perspectives. We are an equal opportunity employer and are committed to providing accommodations throughout the hiring process.

This role uses AI-assisted tools to support initial screening. All assessments and decisions are made by a human reviewer.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead designs and manages GPU infrastructure, Kubernetes deployment, and managed inference platform architecture for Tether's AI compute services.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical lead who designs and manages GPU infrastructure and Kubernetes-based compute platforms for AI inference and containerized workloads.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure, Kubernetes deployment, and platform observability for a distributed compute and inference platform.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Senior Staff Software Engineer, Benefits at Gusto

Senior Staff Software Engineer designs and builds full-stack platform capabilities for benefits applications, focusing on scalable systems and AI-driven interfaces.

Lead Posted 1 day ago RemoteFirstJobs Product
What this role involves

About Gusto

At Gusto, we’re on a mission to grow the small business economy. We handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on their craft and their customers. With teams in Denver, San Francisco, and New York, we support more than 500,000 small businesses nationwide and are building a workplace that reflects the people we serve.

All full-time employees receive competitive base pay, benefits, and equity (RSUs) — because everyone who helps build Gusto should share in its success. Offer amounts are determined by role, level, and location. Learn more about our Total Rewards philosophy.

AI is a fundamental part of how work gets done at Gusto. We expect all team members to actively engage with AI tools relevant to their role and grow their fluency as the technology evolves. AI experience requirements vary by role and will be assessed during the interview process.

About the Role

As a Staff Software Engineer on the Benefits Advisory team, you will be directly accountable for key architectural improvements to Gusto’s benefits platform. This is a full-stack role where you will design and build platform capabilities that enable customers to explore, apply for, and maintain their benefits within your product area.

Your work will focus on transforming our benefits opportunity, shopping and renewal flows, creating clear system boundaries that enable efficient reuse and increased scalability, allowing the ability to offer a wider variety of products more quickly. You will design services that continue to scale the organization, with an emphasis on enabling novel AI-driven interfaces to deliver a delightful benefits shopping experience.

You’ll operate at the intersection of marketing, sales, operations and engineering.  If you are passionate about building highly scalable systems that can reason, predict, and personalize to unlock Gusto Benefits’ next phase of sustainable growth, we’d love to have you join our team.

About the Team

The Benefits Advisory team is building the next generation of infrastructure that powers benefits applications, making it easier than ever to give customers access to the absolute best benefits for their needs.

Our mission is to build a robust system that informs our customers of the most relevant benefits options, making it easier than ever to confidently fulfill benefits that fit our customers current and future needs. We’re building a world-class technology platform designed to understand customer context in real-time, and continuously improve through data and feedback. This means rapid experimentation, fast learning, and strong cross-functional partnership.

We prioritize quality, observability, and uptime because these intelligent systems are fundamental to Gusto’s growth and brand. We partner closely with Marketing, Sales, and Operations to build and connect the AI-powered tools they use every day.

Here’s what you’ll do day-to-day:

  • Architect and evolve our customer-facing web platforms with a focus on scale and performance

Build innovative AI interfaces to best assist our customers

  • Design, build, and maintain shared services to enable rapid iteration on product offerings
  • Implement and optimize key workflows for a best-in-class customer experience.
  • Write high-quality, well-tested code across the stack, leveraging AI-powered tools to accelerate  development and improve reliability
  • Support, mentor, and up-level fellow engineers on the team, helping to establish best practices for AI-native design patterns.
  • Partner cross-functionally with Marketing, Sales, Operations and leadership to translate business needs into technical solutions that fit our needs now and in the future.
  • Influence the technical roadmap for your area of the Benefits Advisory platform, ensuring your work aligns with Gusto’s AI-native strategy and team goals

Here’s what we’re looking for:

  • 8+ years of experience building scalable full-stack web applications with expertise in both frontend (React, TypeScript) and backend (API development, data modeling, database schema design)
  • Extensive Experience building AI interfaces, iterating on prompts and maintaining high quality evals
  • Extensive Experience applying AI tools to accelerate full-stack development, improve code quality, and build intelligent user experiences. Familiar with AI-driven personalization patterns and willing to experiment with AI techniques to improve the benefits shopping experience
  • Deep experience in CI/CD and observability best practices.
  • Ability to act as a thought partner for both technical (Engineering, AI/Data) and business (Sales, Operations) teams.
  • A balance of pragmatic execution and long-term architectural thinking to build a scalable, intelligent platform.
  • Experience enabling productivity across large engineering organizations.
  • Nice to have:
    • Direct familiarity with React, Ruby on Rails
    • Experience with robust recommendation systems
    • AI architecture and scalability for reliable platform services

Compensation

Our cash compensation amount for this role is targeted at $191,000/yr to $225,000/yr in Denver & most remote locations, and $225,000/yr to $265,000/yr for San Francisco, Seattle & New York. Stock equity is additional. Final offer amounts are determined by multiple factors including candidate experience and expertise and may vary from the amounts listed above.

Gusto has physical office spaces in Denver, San Francisco, and New York City. Employees who are based in those locations will be expected to work from the office on designated days approximately 2-3 days per week (or more depending on role). The same office expectations apply to all Symmetry roles, Gusto’s subsidiary, whose physical office is in Scottsdale.

Note: The San Francisco office expectations encompass both the San Francisco and San Jose metro areas.

When approved to work from a location other than a Gusto office, a secure, reliable, and consistent internet connection is required. This includes non-office days for hybrid employees.

Our customers come from all walks of life and so do we. We hire great people from a wide variety of backgrounds, not just because it’s the right thing to do, but because it makes our company stronger. If you share our values and our enthusiasm for small businesses, you will find a home at Gusto.

Gusto is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristic. Gusto considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Gusto is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. We want to see our candidates perform to the best of their ability. If you require a medical or religious accommodation at any time throughout your candidate journey, please fill out this form and a member of our team will get in touch with you.

Gusto takes security and protection of your personal information very seriously. Please review our Fraudulent Activity Disclaimer.

Personal information collected and processed as part of your Gusto application will be subject to Gusto’s Applicant Privacy Notice.

Read the full description
Engineer Staff Software Engineer at Fluxon

Staff engineer who writes production code, guides architectural decisions, mentors teams, and leads technical strategy across AI-driven projects.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Who we are

At Fluxon, we believe that how you build matters as much as what you build. We help businesses navigate their most important technology decisions with confidence, and take responsibility for seeing them through. Founded by ex-Googlers and startup veterans, we’re proud to partner with teams behind some of the most ambitious products, including Google, OpenAI, Anthropic, Walmart and Stripe.

Our work spans strategy, design, and engineering — often in complex, AI-driven environments — where clarity, speed and quality are the standard. We use AI intentionally, applying it only where it adds real value and expands what’s possible. Care shapes everything we do.

Inside Fluxon, you’ll find a global, remote-first team of experienced builders, who are curious, kind and serious about their craft. We’re building a place where people can take ownership, solve problems that matter and do work they’re proud to stand behind. If you want to do your best work alongside people who care as much as you do, you’ll feel at home here.

This role is fully remote, with candidates based in Buenos Aires, Argentina.

About the role

As a Staff Software Engineer at Fluxon, you’ll play a key role in shaping the technical direction of our engineering organization. This is a highly senior, hands-on leadership position where you’ll partner closely with company and engineering leadership to influence strategy, guide architectural decisions, and elevate our overall engineering practice. All Staff Engineers write production code, and everyone joins Fluxon as an individual contributor before stepping into project leadership or management.

You’ll be responsible for:

  • Guiding product delivery all the way to the user, leading projects, providing technical guidance, and building and iterating in a dynamic environment
  • Partnering directly with clients to understand their needs and achieve business goals
  • Defining product requirements, identifying appropriate system designs and planning development in partnership with our Product and Design teams
  • Helping drive a healthy and effective engineering culture within customer teams and inside Fluxon
  • Mentoring engineers across multiple teams, supporting their ongoing growth and strengthening team capabilities

You’ll work with a diversity of technologies, including:

  • Core Languages & Runtimes

    • Primary Languages: TypeScript/JavaScript, Python, Golang (Go), Java, C# (.NET), Kotlin, Swift, Rust
    • Additional: Ruby on Rails, Java, C# (.NET), Kotlin, Swift, Rust
  • Frameworks & Ecosystems

    • Front-End & UI: React, Next.js (Full-Stack), Angular, SwiftUI
    • Back-End & Server-Side: Spring Boot (Java/Kotlin), FastAPI (Python), Django (Python)
    • Mobile Development: Expo (React Native)
  • Cloud & Infrastructure

    • Platforms: Google Cloud Platform (GCP), Amazon Web Services (AWS), Microsoft Azure
    • Compute: GCP Compute Engine (VMs), AWS Fargate (Container Orchestration), Google Cloud Run (Serverless Containers), AWS Amplify
    • Storage: AWS S3 (Object Storage), Google Cloud Storage (GCS)
  • Data & Messaging Services

    • Streaming & Queuing: Apache Kafka (Event Streaming), AWS SQS (Simple Queue Service)
    • Data Warehouse & Analytics: Google BigQuery
    • Monitoring & Observability: GCP Cloud Monitoring Suite (CMS)
  • Data Stores

    • Relational (SQL): PostgreSQL, MariaDB
    • NoSQL/Document: Firestore (Firebase), Supabase, MongoDB
    • In-Memory Caching: Redis, Memcache
  • Advanced Technologies & Architecture

    • Artificial Intelligence/Machine Learning: AI/ML, Large Language Models (LLMs), Agentic AI, Natural Language Processing (NLP)
    • LLM Platforms: Google Gemini, OpenAI ChatGPT, Vertex AI (GCP), Anthropic Claude, Hugging Face (OSS Models)
    • Software Design: Architecture Redesign, Single Page Applications (SPA), Mobile Application Development
    • Emerging Tech: Blockchain/Crypto

Qualifications

  • 7+ years of industry experience in software development
  • Experience leading development through the full product lifecycle, including CI/CD, testing, release management, deployment, monitoring and incident response
  • Fluent in the design and implementation of scalable system architectures, data structures and algorithms, and effective development practices
  • Professional fluency in written and spoken English, with the ability to communicate technical concepts clearly to both international teammates and customers

What we offer

  • Remote-first, flexible work with a budget to set up a work space that works for you
  • Localized healthcare coverage to support you and your family’s wellbeing
  • Flexible paid time off with a minimum of 3 weeks per year (plus holidays)
  • “No internal meetings” Fridays, so you can focus on deep, uninterrupted work
  • A professional growth budget for learning that matters to you –  whether it’s developing your technical skills or learning a new language
  • A monthly wellness allowance to support your physical and mental health
  • Annual company offsites, where we gather in person to build meaningful connections
  • A competitive salary that reflects your expertise, impact and experience
  • Profit-sharing, so you can benefit directly from the value we build together
  • A paid sabbatical program, designed for rest, renewal and fresh perspective

We believe diverse teams perform better, and an inclusive environment is essential to building a successful organization. We welcome applicants from all backgrounds, experiences, and perspectives. We are an equal opportunity employer and are committed to providing accommodations throughout the hiring process.

This role uses AI-assisted tools to support initial screening. All assessments and decisions are made by a human reviewer.

Read the full description
Engineer Lead Data Engineer at Box

Designs and manages scalable data platforms, pipelines, and governance frameworks while leading cross-functional teams on data architecture and AI-driven initiatives.

Lead Posted 1 day ago RemoteFirstJobs Product
What this role involves

\*\*\* This is where your organization can create a consistent intro to all of your jobs, creating consistency in voice and messaging across all job posts

\*\*\* C’est ici que votre organisation peut crĂ©er une introduction cohĂ©rente Ă  tous vos emplois, en crĂ©ant une cohĂ©rence dans la voix et la messagerie dans tous les postes.

Marlabs, a global AI and Digital Solutions Consulting firm, delivers intelligent solutions across AI, data, analytics, and product engineering. Since 2000, we have partnered with some of the largest healthcare, life sciences, financial services, and government organizations worldwide. As we continue to expand our global footprint, we have an exciting opportunity for a highly skilled Lead Data Engineer to join our innovative and dynamic team.

Lead Data Engineer | About You

As a Lead Data Engineer, you will be responsible for designing, building, and managing the organization’s modern data platform, ensuring reliable, secure, and scalable data products that support analytics, reporting, and AI-driven business initiatives. You will lead the development of enterprise data pipelines, lakehouse architecture, and governance frameworks while partnering closely with AI/ML, platform engineering, and security teams. The ideal candidate combines deep expertise in data engineering, data modeling, cloud-based architectures, and data governance with a strong focus on reliability, observability, and regulatory compliance.

Lead Data Engineer | Day-to-Day

  • Design, build, and maintain scalable lakehouse architectures and enterprise data platforms that support analytics, reporting, and AI-driven solutions.
  • Develop and manage secure data ingestion frameworks, including CDC, batch, API, and file-based integrations from operational and transactional source systems.
  • Create and maintain data models, semantic layers, and governed metrics that enable consistent, trusted, and business-ready data consumption.
  • Implement data quality, reconciliation, observability, and monitoring processes to ensure reliable, recoverable, and high-performing data pipelines.
  • Partner with AI/ML, platform engineering, and security teams to deliver governed data products, support regulatory compliance requirements, and ensure proper data classification and access controls.
  • Establish and enforce data governance standards, lineage documentation, data contracts, quality thresholds, and operational procedures while supporting onboarding of new data sources and environments.

Lead Data Engineer | Skills & Experience

  • 8+ years of experience in Data Engineering, including end-to-end ownership of data ingestion, transformation, storage, and analytics delivery, with experience leading large-scale data initiatives.
  • Strong expertise in Python and SQL, including advanced data modeling, transformation frameworks, data quality management, and performance optimization.
  • Hands-on experience with modern lakehouse architectures and data platforms, including technologies such as Apache Iceberg, Delta Lake, Hudi, object storage, and cloud-native data solutions.
  • Proven experience building scalable data pipelines and CDC solutions, leveraging technologies such as Kafka, Debezium, Airflow, Dagster, dbt, and enterprise integration frameworks.
  • Strong understanding of data governance, lineage, security, and compliance practices, including data contracts, access controls, observability, audit readiness, and regulated industry environments.
  • Experience collaborating with AI/ML, analytics, platform engineering, and business teams to deliver trusted, governed, and scalable data products; financial services or banking industry experience is highly preferred.

\*\*\* Similar to the introduction that can precede all job descriptions, an outro can also be formatted for consistency on all posts

\*\*\* Semblable Ă  l’introduction qui peut prĂ©cĂ©der toutes les descriptions de poste, une outro peut Ă©galement ĂȘtre formatĂ©e pour la cohĂ©rence sur tous les messages

Read the full description
Engineer Staff Software Engineer at Fluxon

Staff-level software engineer who guides technical strategy, leads projects, mentors engineers, and writes production code across AI-driven and complex systems.

Lead Remote Posted 1 day ago RemoteFirstJobs Product
What this role involves

Who we are

At Fluxon, we believe that how you build matters as much as what you build. We help businesses navigate their most important technology decisions with confidence, and take responsibility for seeing them through. Founded by ex-Googlers and startup veterans, we’re proud to partner with teams behind some of the most ambitious products, including Google, OpenAI, Anthropic, Walmart and Stripe.

Our work spans strategy, design, and engineering — often in complex, AI-driven environments — where clarity, speed and quality are the standard. We use AI intentionally, applying it only where it adds real value and expands what’s possible. Care shapes everything we do.

Inside Fluxon, you’ll find a global, remote-first team of experienced builders, who are curious, kind and serious about their craft. We’re building a place where people can take ownership, solve problems that matter and do work they’re proud to stand behind. If you want to do your best work alongside people who care as much as you do, you’ll feel at home here.

This role is fully remote, with candidates based in Buenos Aires, Argentina.

About the role

As a Staff Software Engineer at Fluxon, you’ll play a key role in shaping the technical direction of our engineering organization. This is a highly senior, hands-on leadership position where you’ll partner closely with company and engineering leadership to influence strategy, guide architectural decisions, and elevate our overall engineering practice. All Staff Engineers write production code, and everyone joins Fluxon as an individual contributor before stepping into project leadership or management.

You’ll be responsible for:

  • Guiding product delivery all the way to the user, leading projects, providing technical guidance, and building and iterating in a dynamic environment
  • Partnering directly with clients to understand their needs and achieve business goals
  • Defining product requirements, identifying appropriate system designs and planning development in partnership with our Product and Design teams
  • Helping drive a healthy and effective engineering culture within customer teams and inside Fluxon
  • Mentoring engineers across multiple teams, supporting their ongoing growth and strengthening team capabilities

You’ll work with a diversity of technologies, including:

  • Core Languages & Runtimes

    • Primary Languages: TypeScript/JavaScript, Python, Golang (Go), Java, C# (.NET), Kotlin, Swift, Rust
    • Additional: Ruby on Rails, Java, C# (.NET), Kotlin, Swift, Rust
  • Frameworks & Ecosystems

    • Front-End & UI: React, Next.js (Full-Stack), Angular, SwiftUI
    • Back-End & Server-Side: Spring Boot (Java/Kotlin), FastAPI (Python), Django (Python)
    • Mobile Development: Expo (React Native)
  • Cloud & Infrastructure

    • Platforms: Google Cloud Platform (GCP), Amazon Web Services (AWS), Microsoft Azure
    • Compute: GCP Compute Engine (VMs), AWS Fargate (Container Orchestration), Google Cloud Run (Serverless Containers), AWS Amplify
    • Storage: AWS S3 (Object Storage), Google Cloud Storage (GCS)
  • Data & Messaging Services

    • Streaming & Queuing: Apache Kafka (Event Streaming), AWS SQS (Simple Queue Service)
    • Data Warehouse & Analytics: Google BigQuery
    • Monitoring & Observability: GCP Cloud Monitoring Suite (CMS)
  • Data Stores

    • Relational (SQL): PostgreSQL, MariaDB
    • NoSQL/Document: Firestore (Firebase), Supabase, MongoDB
    • In-Memory Caching: Redis, Memcache
  • Advanced Technologies & Architecture

    • Artificial Intelligence/Machine Learning: AI/ML, Large Language Models (LLMs), Agentic AI, Natural Language Processing (NLP)
    • LLM Platforms: Google Gemini, OpenAI ChatGPT, Vertex AI (GCP), Anthropic Claude, Hugging Face (OSS Models)
    • Software Design: Architecture Redesign, Single Page Applications (SPA), Mobile Application Development
    • Emerging Tech: Blockchain/Crypto

Qualifications

  • 7+ years of industry experience in software development
  • Experience leading development through the full product lifecycle, including CI/CD, testing, release management, deployment, monitoring and incident response
  • Fluent in the design and implementation of scalable system architectures, data structures and algorithms, and effective development practices
  • Professional fluency in written and spoken English, with the ability to communicate technical concepts clearly to both international teammates and customers

What we offer

  • Remote-first, flexible work with a budget to set up a work space that works for you
  • Localized healthcare coverage to support you and your family’s wellbeing
  • Flexible paid time off with a minimum of 3 weeks per year (plus holidays)
  • “No internal meetings” Fridays, so you can focus on deep, uninterrupted work
  • A professional growth budget for learning that matters to you –  whether it’s developing your technical skills or learning a new language
  • A monthly wellness allowance to support your physical and mental health
  • Annual company offsites, where we gather in person to build meaningful connections
  • A competitive salary that reflects your expertise, impact and experience
  • Profit-sharing, so you can benefit directly from the value we build together
  • A paid sabbatical program, designed for rest, renewal and fresh perspective

We believe diverse teams perform better, and an inclusive environment is essential to building a successful organization. We welcome applicants from all backgrounds, experiences, and perspectives. We are an equal opportunity employer and are committed to providing accommodations throughout the hiring process.

This role uses AI-assisted tools to support initial screening. All assessments and decisions are made by a human reviewer.

Read the full description
Engineer Staff Software Engineer, PostgreSQL at Beacon Biosignals

Design and scale backend data infrastructure systems, PostgreSQL schemas, and event pipelines for a clinical brain data platform.

Lead Posted 1 day ago RemoteFirstJobs Product
What this role involves

Beacon Biosignals is transforming precision medicine for the brain, from clinical development to clinical care. For Life Sciences partners, we offer the leading at-home EEG platform for clinical development of novel therapeutics for neurological, psychiatric, and sleep disorders. Our Diagnostics business is building the most comprehensive at-home platform for precision diagnostics, combining EEG and cardiopulmonary signals to deliver reimbursable assessments for sleep and central nervous system disorders. Together, we’re changing the way patients are diagnosed and treated for any disorder that affects brain physiology.

Beacon Biosignals is seeking a Software Engineer IV to join our Datastore team. The Datastore sits at the heart of Beacon’s platform. It defines our foundational data model and serves as the centralized repository and API for brain data collected from clinical studies, supporting the analytical tools, services, and applications used by our scientists, clinicians, and partners.

This role emphasizes the design and scaling of backend systems for data infrastructure: PostgreSQL data models, Kafka-driven event pipelines, and the GraphQL API that makes scientific and clinical data usable - distinct from frontend development or end-user product surfaces. At this level you will lead the design of complex systems within the Datastore rather than individual components, own the pipelines that deliver dataset snapshots into our warehouse, and collaborate across platform, application and scientific teams to implement APIs that unlock new workflows in Beacon’s expanding portfolio of clinical studies and digital health products.

You should have experience writing complex SQL queries by hand, and should be prepared to reason about schema design, indexes, and query behavior as a matter of course.

Two properties of our system that shape our work may influence your application. Our GraphQL API is generated from the PostgreSQL schema using PostGraphile, so a schema decision is an API decision. Secondly, row-level security is in force on effectively every table, because clinical data from different studies and partners shares one schema, with some rows shared for access by multiple internal teams. RLS predicates get injected into your queries, which makes reasoning about query plans genuinely harder here than it is in most places. If that sounds interesting rather than tedious, you’ll enjoy this team.

We are hiring two engineers into this role, with two different centers of gravity. One will focus on the breadth of the Datastore: the event pipelines, the GraphQL API, core data models, and third-party integrations. The other will focus on depth in PostgreSQL (data modeling, performance tuning, operations) and supporting all engineers in developing schemas and the access controls that protect clinical data. The core requirements are the same. Tell us which one you are most drawn to in your application.

This role is 100% remote from anywhere in the U.S. Beacon’s robust asynchronous work practices ensure a first-class remote work experience, but we also have in-person office hubs in Boston, New York City, and Paris.

What success looks like

  • You lead design and architecture for complex projects and systems within the Datastore, weighing implementation trade-offs with engineering principles and backing your decisions with data rather than preference
  • You partner with product managers, scientific and clinical stakeholders, and other development and support teams to translate their requirements into new integrations, services, and solutions for scaling Beacon’s data platform - contributing to broader product and technical decisions, not simply implementing them
  • You debug and profile ambiguous problems that cross system boundaries, including codebases beyond the Datastore’s purview, and you leave what you touch better than you found it — structure, test coverage, tooling
  • Peers look to your work as an example: you implement features with attention to detail, demonstrating product awareness and high standards for data integrity, design soundness, and maintainability across the lifespan of the system
  • You draft RFCs to lead complex feature discovery and technical design phases, soliciting and incorporating input, and helping your team align on a set of decisions to guide development
  • You plan your own work and contribute to planning the team’s, making reasoned trade-offs between speed, generalizability, and technical debt, and accounting for the needs of multiple stakeholder teams
  • You teach and mentor other engineers regularly, give feedback that people act on, and are sought out for it. You may act as tech lead on a project or as the coordination point for engineering practice within the team
  • You leverage feedback from internal stakeholders and external partners to improve operational robustness and ease of use of the Datastore and its services via new documentation, tools, and dashboards

What you will bring

  • 7+ years of backend development experience, including at least 3 years focused on data platforms and backend infrastructure for data-centric applications: transactional systems, event pipelines, and the APIs that serve them
  • Highly proficient in at least one technical area relevant to this work: data modeling, distributed systems, API design, or database engineering, with the depth that others on a team come to you for
  • Strong proficiency in SQL, with production experience working with PostgreSQL as an application database; you write and tune queries fluently and are comfortable reasoning about schema design, indexes, and query plans
  • Hands-on experience with data streaming and event processing, using Kafka (our event bus) or an equivalent such as RabbitMQ, Pulsar, or a cloud-based equivalent like Amazon Kinesis
  • Proficiency in Julia or Python; Julia is a core Datastore language, and experience with Julia is a strong plus, but not required — we welcome candidates eager to learn it!
  • Experience deploying and operating services in containerized environments, with familiarity in Kubernetes and Infrastructure as Code (IaC) tools such as Terraform and/or Helm
  • A track record of mentoring engineers and improving how a team works - practices, tooling, test coverage, or documentation that outlasted your involvement
  • A collaborative mindset: you enjoy working across disciplines and believe people achieve more together than alone
  • Excellent written and verbal communication skills, especially in asynchronous and remote-friendly environments
  • A self-directed approach, with a track record of thriving in hybrid or fully remote teams
  • Experience with, or interest in, using LLM-assisted or agentic coding tools in a production setting, with good judgment about what guardrails are needed

For the Datastore Systems role:

  • You can lead design across the breadth of the platform.
  • GraphQL and TypeScript/JavaScript (Node.js) - GraphQL is the Datastore’s primary interface, and this profile owns significant parts of it
  • Designing event-driven pipelines end to end: producers, consumers, replay, idempotency, and what happens when a consumer falls behind
  • Leading the design of complex systems rather than individual components, across the API, the pipelines, and the services that depend on them

For the PostgreSQL & Data Modeling role:

  • You can be the person the rest of the team comes to about the database.
  • Deep PostgreSQL: you reason fluently about the query planner, index selection, and EXPLAIN (ANALYZE), and why a query that should be fast isn’t
  • Data modeling: you design the schema itself: entities, relationships, and the constraints that make invalid states impossible. Including how to model time, since clinical results get re-scored and corrected and we need to answer “what did this look like when the report was signed?”
  • Schema evolution under load: zero-downtime migrations, dual-write transitions, lock avoidance, backfill and verification. Our schema has hundreds of migrations behind it and will have hundreds more
  • Operating PostgreSQL in production, not only designing for it: upgrades, replication, backup and restore, connection pooling, bloat and vacuum behaviour, extensions
  • Making others good at it: you teach, review, and document, because a schema only stays coherent if the team understands it

We don’t expect every candidate to check every box. If you bring strong fundamentals, care about data infrastructure for health, and are excited to learn, we’d love to hear from you!

The US-based salary range for this role is $170,000 – $190,000. Salary ranges are determined using current market compensation data for this role and adjusted based on experience, skills, and location. The base salary is one component of the total compensation package, which includes equity, PTO and other benefits.

At Beacon, we’ve found that cultural and scientific impact is driven most by those that lead by example. As such, we’re always seeking new contributors whose work demonstrates an avid curiosity, a bias towards simplicity, an eye for composability, a self-service mindset, and - most of all - a deep empathy towards colleagues, stakeholders, users, and patients. We believe a diverse team builds more robust systems and achieves higher impact.

Read the full description
Engineer Track it Forward: Lead Developer — Rebuild, Modernize, & Scale (Social Good SaaS, Remote)

Lead developer rebuilds and modernizes a legacy SaaS platform, choosing tech stack, architecting systems, and migrating customers incrementally while scaling operations.

Lead Remote Posted 2 days ago We Work Remotely — Programming
What this role involves

Headquarters: Oakland, CA
URL: https://www.trackitforward.com

We are looking for a lead developer to do a greenfield rebuild of our legacy SaaS app, modernize our mobile apps, and lead the tech effort to scale our app.  You will help choose the language, build the technical foundation, work with our product manager to rebuild the app, and migrate all customers over.

About Track it Forward Track it Forward is the leading volunteer time-tracking application for nonprofits and schools. For 15 years, we’ve helped thousands of organizations mobilize volunteers and track over 30 million hours. We are a profitable, bootstrapped team of 4 focused on great work-life balance, high-impact work, and a fun, puzzle-loving culture (daily crosswords, monthly virtual escape rooms, and annual retreats).  We are not a VC-funded startup. We grew organically, stayed profitable, and prioritize high-quality work and personal fulfillment without growth-at-all-costs pressure. The team is led by the original founder and developer.

The Modernization Strategy This is a structured, phased "small bang" migration—not a high-risk cutover or a strangler-fig migration. You will build a parallel greenfield system, extract legacy components step-by-step, and migrate customers over incrementally tenant-by-tenant while keeping operations running smoothly.

  • Current Tech: Drupal 6 architecture hosted on Pantheon, paired with two Ionic/Capacitor mobile apps.

  • Your Scope: Complete ownership of the technical architecture. You will partner with our product manager, establish engineering processes, and lead the rebuild from scratch.

  • Modern AI practices:  We want you to be on the cusp of AI workflows while also balancing human guardrails to produce quality code.

  • Team Structure: You will be our sole developer in the short term, taking full technical helm as we scale.

Target Stack & Responsibilities

  • Core Tech Stack: Core backend/frontend rebuild focused on Python/Django, Laravel, or React. You will make the final architectural call with input from the CEO/founder.

  • Database & Storage Extraction: Deconstruct our legacy Drupal 6 / MySQL schema into a clean, domain-driven PostgreSQL database, and offload 100GB+ of media files to AWS S3 or Cloudflare R2.

  • Security & Isolation: Implement application-layer field-level encryption (AES-256 for PII, Bcrypt/Argon2id for passwords) and build zero-trust dev environments so contractors never handle production user data.

  • CI/CD & Documentation: Establish automated testing pipelines, clean API contracts, and clear runbooks so the codebase remains maintainable long-term.

  • Mobile App Maintenance: Maintain and update our two Ionic/Capacitor mobile apps using AI coding tools. Mobile is a small fraction of the overall role—you just need to leverage AI to keep them functional without being a native mobile expert.

Candidate Qualifications

  • Pre-AI Engineering Foundation: 5+ years of core software engineering and architectural experience established before generative AI existed. You hold strong first-principles opinions on database normalization and system patterns.

  • Proven Migration Experience: You have successfully architected and led at least one legacy monolith migration. You know what inputs and scope decisions keep momentum going.

  • Modernized AI Velocity: You actively use AI tools to multiply output, but retain strict architectural control, unit testing discipline, and manual code review standards.

Location & Remote Flexibility

  • US W-2 Tax Residency Required: You must be legally authorized to work in the US and maintain primary US tax residency (for payroll, contracts, and legal/tax compliance).

  • Travel-Friendly: We are 100% remote. As long as your primary legal/tax home remains in the US, you are welcome to travel or digital nomad—provided you can consistently hit our core sync hours.

  • Core Sync Hours: Mandatory availability from 9:00 AM – 12:00 PM PST for daily team syncs and real-time collaboration. Flexible hours outside this window.

Compensation & Benefits

  • Base Salary: $150,000 – $200,000 / year (commensurate with experience)

  • Benefits: 401(k) w/ company match, Health Insurance, vacation, PTO, and holdiays

  • Culture: High autonomy, direct project ownership, and zero venture-capital growth pressure

How to Apply To filter out automated spam: include the most authentically human, non-AI reason why this specific role resonates with you.  Specifically include a short Loom video of you in the “Why should we choose you over any one else?” question.  This isn’t required but will sure stand out.  Be authentic, no need to do crazy rehearsals, just be you.

Apply at this google form here

 

To apply: https://weworkremotely.com/remote-jobs/track-it-forward-lead-developer-rebuild-modernize-scale-social-good-saas-remote

Read the full description
Engineer Member of the Technical Staff - Data Platform at Vercel

Staff engineer designs and leads development of a next-generation data platform supporting real-time analytics, ETL processes, and ML initiatives using Kafka, ClickHouse, and Snowflake.

Lead Posted 3 days ago RemoteFirstJobs Product
What this role involves

About Vercel:

Vercel is the agentic infrastructure company. We free people and agents to ship what’s next.

For more than a decade, Vercel has shaped how the web is built. As the team behind Next.js, v0, and AI SDK, we create products that help builders move from idea to production with speed, security, and exceptional developer experience.

Now, software is entering a new era, and the next generation of products will not just be used by people. They will be built, extended, and operated by agents.

We are building the platform for that future, trusted by companies like OpenAI, PayPal, Ramp, Supreme, and millions of developers worldwide. Whether you’re building our products, supporting our customers, growing our community, or shaping our story, you’ll help define what comes next.

About the Role

We’re looking for a Staff level Engineer to lead the design and development of a next-generation Data Platform at Vercel. This is a unique opportunity to architect a foundational system that will power data across our products and business, supporting everything from real-time analytics to AI/ML initiatives.

Reporting directly to the Senior Director of Data Engineering, you’ll define the technical vision for our data ecosystem, work across the organization to align engineering strategy with product goals, and build scalable, reliable systems using modern technologies like Kafka, ClickHouse, Tinybird, and Snowflake.

You’ll lead from the front, writing production-grade code, setting high standards for design and data governance, and mentoring a team of engineers as the organization grows. If you’re excited by the challenge of building scalable data systems that drive product innovation, we’d love to hear from you.

What You Will Do

  • Design and implement a next-generation data platform that supports diverse requirements, including batch and real-time integrations, advanced analytics, and multiple data types, leveraging Kafka, Kafka-based tooling, and additional streaming technologies to power real-time data movement and processing.
  • Work with engineering, product, and senior leadership teams across Vercel to architect solutions that address business needs, aligning technical strategies with product goals and ensuring seamless data availability and usability.
  • Define and maintain guidelines for data ingestion, creation, enrichment, and storage, while overseeing data modeling, ETL processes, and data warehousing (ClickHouse, Tinybird, Snowflake) to achieve low-latency analytics and scalable storage.
  • Develop end-to-end solutions across the tech stack, setting the example for engineering excellence, and provide mentorship and direction that fosters continuous improvement and innovation in the team.
  • Partner with Security, Compliance, and Legal teams to ensure adherence to data protection standards, and champion high availability and fault tolerance through comprehensive monitoring, alerting, and incident response practices.
  • Drive architectural decisions and roadmap planning that ensure efficiency, scalability, and high performance, and evaluate build-vs-buy trade-offs, integrating vendor solutions where appropriate.
  • Collaborate with data scientists and machine learning teams to evolve infrastructure that supports advanced analytics and AI initiatives, serving as a key contributor to strategy, tooling, and platform enhancements for ML-related workloads.
  • Act as a key decision-maker for discovering, vetting, and selecting data architectures and technologies critical to Vercel’s success, and establish multi-phase strategic roadmaps that align with overall business objectives and drive data platform maturity.

About You

  • 8+ years of experience in data engineering, data architecture, or related roles, with at least 5 years at the Principal Engineer level.
  • Proven track record designing and operating large-scale data infrastructures at high scale in a complex, fast-paced environment, with experience across Kafka and its ecosystem (e.g., Kafka Streams, Confluent Platform), ClickHouse, Tinybird, Snowflake, and broader big data frameworks.
  • Proficiency in cloud platforms (AWS, GCP, or Azure) and associated big data services.
  • Strong background in data governance and security, and adept at ensuring compliance with regulatory standards and protecting sensitive information.
  • Outstanding communication and collaboration skills, with the ability to influence technical and non-technical stakeholders across varying organizational levels.
  • A leadership mindset, with demonstrated ability to mentor teams, drive consensus, and advocate for best practices across the organization.
  • A Master’s degree in Computer Science, Engineering, or a related field is preferred.
  • Industry recognition or notable contributions in data engineering (e.g., published works, open-source contributions) is a plus.

The San Francisco, CA base pay range for this role is $232,000 - $348,000. Actual salary will be based on job-related skills, experience, and location. Compensation outside of San Francisco may be adjusted based on employee location. The total compensation package may include benefits, equity-based compensation, and eligibility for a company bonus or variable pay program depending on the role. Your recruiter can share more details during the hiring process.

Vercel is committed to fostering and empowering an inclusive community within our organization. We do not discriminate on the basis of race, religion, color, gender expression or identity, sexual orientation, national origin, citizenship, age, marital status, veteran status, disability status, or any other characteristic protected by law. Vercel encourages everyone to apply for our available positions, even if they don’t necessarily check every box on the job description.

Read the full description
Engineer Staff Machine Learning Engineer - Retention at Taskrabbit

Staff ML Engineer who develops and owns machine learning models for customer retention, tasker matching, and repeat purchase optimization at scale.

Lead Hybrid Posted 3 days ago RemoteFirstJobs Product
What this role involves

About Taskrabbit:

Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more.

At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world.

Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed!

This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday.

About the Role

Machine Learning is a cornerstone at Taskrabbit, and we’re looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale.

Taskrabbit’s greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There’s significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expansion unlocks new revenue streams within our existing customer base.

This role is central to capturing that opportunity. While initial matching quality and service discovery matter, the real competitive advantage and growth lever is optimizing the experience after a successful first job.

What You’ll Work On:

  • Taskrabbit Ranking Model: Own the reliability and performance of our core ranking system, ensuring accurate tasker-to-job matching and optimizing First-Time Right (FTR) rates.
  • Increased repeat purchase frequency through intelligent matching, personalized recommendations, and category discovery
  • Expanded customer lifetime value by helping customers find and return for new service categories
  • Optimized affordability and relevance via dynamic pricing, smart segmentation, and category-specific experiences
  • Reduced friction and churn through predictive quality interventions and proactive customer success
  • Marketplace resilience by building systems that keep high-value customers engaged and loyal
  • End-to-End ML Lifecycle: Own the complete lifecycle of models—from feature engineering and training through evaluation, deployment, monitoring, and optimization in production.
  • Infrastructure & Scalability: Build and maintain scalable, reliable ML infrastructure and data pipelines that support reproducible feature engineering and model deployment across real-time, near real-time, and batch contexts.
  • Monitoring & Performance Optimization: Develop monitoring and observability systems to understand data quality and model performance in complex systems. Collaborate with engineering and science teams to optimize algorithms for training, inference, and evaluation.
  • Software Engineering Excellence: Write clean, efficient, and maintainable code. Participate actively in code reviews, documentation, and best practices across the full software engineering lifecycle.

Your Areas of Expertise:

We welcome applicants from a variety of backgrounds and experiences. Below gives you a sense of how we’re thinking about what you’ll need to be successful in the role.

  • BS, MS, or PhD in Computer Science, Statistics, Operations Research, or a related quantitative field.
  • 8+ years of industry experience building and deploying high-quality, production-grade machine learning models and systems.
  • Strong theoretical knowledge and hands-on experience in machine learning, particularly in search, ranking, recommender systems, pricing/elasticity modeling, or predictive analytics.
  • Solid software engineering skills with proficiency in one or more programming languages, including Python.​​  The candidate should have experience with popular ML libraries like Scikit-learn, lightgbm, xgboost, TensorFlow, PyTorch, etc.
  • Proficiency in SQL is also required for writing complex queries and transforming data.
  • Experience building REST API-based services.
  • Experience with modern data and ML technologies, such as Docker, Kubernetes, Kafka, Airflow, data warehouses (eg snowflake, redshift or BigQuery), and data lakes.
  • Familiarity with dbt is a plus for transforming and testing data.
  • Familiarity with tools for Infrastructure as Code, such as Github actions, and CI/CD pipelines.
  • Excellent communication skills, with the ability to present complex findings and recommendations clearly to both technical and non-technical audiences.
  • A passion for quickly learning new technologies and a drive to solve challenging problems, and a collaborative mindset.
  • Ideally, experience working in marketplace or platform contexts where ranking, matching, and pricing directly impact user experience and business outcomes.

Compensation & Benefits:

At Taskrabbit, our approach to compensation is designed to be competitive, transparent, and equitable. Total compensation consists of base pay + bonus + benefits + perks. The base pay range for this position is $170,000 - $225,000. This range is representative of base pay only, and does not include any other total cash compensation amounts, such as company bonus or benefits. Final offer amounts may vary from the amounts listed above and will be determined by factors including, but not limited to, relevant experience, qualifications, geography, and level.

You’ll love working here because:

  • Taskrabbit is a Hybrid Company. We value flexibility and choice but also stay committed to regular in-person connection.
  • The People. You will be surrounded by some of the most talented, supportive, smart, and kind leaders and teams – people you can be proud to work with!
  • The Diverse Culture. We believe that we make better decisions when our workforce reflects the diversity of the communities in which we operate. Women make up half of our leadership team and our diversity representation is above that of the tech industry average.
  • The Perks. Taskrabbit offers our employees with employer-paid health insurance and a 401k match with immediate vesting for our US based employees. We offer all of our global employees generous and flexible time off with 2 company-wide closure weeks, Taskrabbit product stipends, wellness + productivity + education stipends, IKEA discounts, reproductive health support, and more. Benefits vary by country of employment.

Taskrabbit’s commitment to Diversity and Inclusion:

An Active Commitment to Equity within our Company and Platform. We are an inclusive community where all who share our mission and values belong. Our diverse team represents the communities we serve, breaking down systemic barriers, and transforming lives- one action at a time.

Taskrabbit is an equal opportunity employer and values diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, ancestry, citizenship, sex, gender, gender identity, sexual orientation, age, marital status, military/veteran status, or disability status. Taskrabbit is committed to working with and providing reasonable accommodation to applicants with physical and mental disabilities. Taskrabbit will consider for employment all qualified applicants with criminal histories in a manner consistent with applicable law.

AI-Assisted Prescreening Notice [US Based Candidates Only]: As part of our hiring process, we may use artificial intelligence tools to assist with the initial prescreening of applications and responses. This tool does not make hiring decisions — every application and response is reviewed by a member of our recruiting team to determine fit for the role. If you would prefer not to have your application processed using this tool, you may opt out by selecting ‘opt out’ in the application form or by emailing talentacquisition@taskrabbit.com, and your application will be reviewed manually with no impact on your candidacy.

Read the full description