Skip to main content

Fixing Oozie LZO ClassNotFound Errors and ShareLib Issues

Struggling with delivery, architecture alignment, or platform stability?

I help teams fix systemic engineering issues: processes, architecture, and clarity.
→ See how I work with teams.


When LZO compression is enabled in Hadoop, Oozie may fail to launch MapReduce jobs with a ClassNotFoundException for the LzoCodec. This article explains why it happens, how to fix the missing hadoop-lzo classes on the Oozie server, how to correctly deploy the Oozie sharelib and how to enable the legacy uber-jar feature for MapReduce actions in CDH-era clusters.

Note (2025): This article documents behaviour from the CDH 4.x/5.x Oozie + MapReduce stack. Oozie is legacy and many teams migrate to Apache Airflow, Dagster, Argo or cloud-native schedulers, but troubleshooting old clusters still requires understanding these patterns. The solutions here apply to Hadoop MR/LZO deployments still in maintenance or migration mode.

Symptom: Oozie fails after enabling LZO compression

When you add LZO codecs to core-site.xml, Oozie may suddenly fail to start any MapReduce job. Typical configuration looks like:

<property>
  <name>io.compression.codecs</name>
  <value>
    org.apache.hadoop.io.compress.GzipCodec,
    org.apache.hadoop.io.compress.DefaultCodec,
    com.hadoop.compression.lzo.LzoCodec,
    com.hadoop.compression.lzo.LzopCodec,
    org.apache.hadoop.io.compress.BZip2Codec
  </value>
</property>

<property>
  <name>io.compression.codec.lzo.class</name>
  <value>com.hadoop.compression.lzo.LzoCodec</value>
</property>

After restarting services, Oozie reports:

java.lang.ClassNotFoundException:
  Class com.hadoop.compression.lzo.LzoCodec not found

Root cause

Even though LZO is installed on Hadoop nodes, Oozie runs inside its own server JVM and classpath. If hadoop-lzo.jar is missing from /var/lib/oozie/ (or the directory your Oozie installation loads from), Oozie cannot load the LZO codec classes and refuses to start jobs.

Fix: Install the LZO JAR on the Oozie server

Copy (or symlink) the hadoop-lzo.jar into the Oozie library directory and restart Oozie:

cp /usr/lib/hadoop/lib/hadoop-lzo.jar /var/lib/oozie/
service oozie restart

Once Oozie can load com.hadoop.compression.lzo.LzoCodec, MapReduce jobs using LZO compression start normally.

The second common issue: missing or incorrect ShareLib

Even with the LZO JAR fixed, Oozie may still fail if the sharelib is missing or has wrong permissions. The sharelib provides the job launcher classpath for Oozie workflows.

Typical setup:

# Create Oozie home in HDFS
sudo -u hdfs hadoop fs -mkdir /user/oozie
sudo -u hdfs hadoop fs -chown oozie:oozie /user/oozie

# Extract the sharelib
mkdir /tmp/share
cd /tmp/share
tar xvfz /usr/lib/oozie/oozie-sharelib.tar.gz

# Upload it to HDFS
sudo -u oozie hadoop fs -put share /user/oozie/share

After uploading, restart Oozie or run:

oozie admin -sharelibupdate

Without a valid sharelib, Oozie cannot assemble the runtime classpath for MapReduce actions and will fail even if the LZO JAR is present on the Oozie server.

Uber JAR support (CDH 4.1+)

Starting with CDH 4.1, Oozie introduced a lightweight uber-jar feature. An uber-jar doesn’t bundle all dependencies into one fat JAR; instead, it carries references to dependent libraries in an internal lib/ directory.

To enable it globally, set the following in oozie-site.xml:

<property>
  <name>oozie.action.mapreduce.uber.jar.enable</name>
  <value>true</value>
</property>

Once enabled, workflow authors can tag a MapReduce action with:

oozie.mapreduce.uber.jar = <path-to-jar>

This tells Oozie that the provided JAR should be treated as an uber-jar, and Oozie will expand and distribute its libraries accordingly when launching the job.

Summary

  • LZO in core-site.xml often breaks Oozie because Oozie’s server classpath cannot see hadoop-lzo.jar.
  • Fix by copying hadoop-lzo.jar into /var/lib/oozie/.
  • Ensure /user/oozie/share exists and contains a valid sharelib.
  • Enable uber-jar support if your workflows depend on it.

With these steps, Oozie resumes launching MapReduce jobs correctly even when LZO compression is enabled cluster-wide.

If you need help with distributed systems, backend engineering, or data platforms, check my Services.

Most read articles

Building a Model-Agnostic Multi-Agent System with OpenClaw

Over one week we rebuilt our AI stack around OpenClaw’s multi-agent architecture to avoid provider lock-in and stop wasting premium tokens. By aligning models to tasks, diversifying fallbacks across providers, enforcing minimal tool access, and switching to memory-first workflows with ephemeral sessions, we reduced token usage per task by about 70% and cut our monthly bill by 77% while improving operational resilience. How We Achieved 77% Cost Reduction and Provider Independence Over the past week, we rebuilt our AI infrastructure around OpenClaw’s multi-agent architecture. The result was a 77% cost reduction , provider independence , and a delegation system that routes work to the most cost-effective model for each job. Below is the technical journey of optimizing a 7-agent squad with OpenClaw. The Challenge: Model Provider Lock-In We started with a simple problem: our entire squad defaulted to a single model provider. This created three issues: Cost inefficiency beca...

BacNet => MQTT in Production: The Real Cost of Bridging BACnet to MQTT at Scale

bacnet2mqtt looks simple in a README and expensive in production. Once BACnet polling, reconnection behavior, stale state, and MQTT publishing collide, teams discover they are not deploying a lightweight adapter but operating infrastructure. This article breaks down where bacnet2mqtt works, where it becomes a bottleneck, and which production patterns reduce the operational damage before incidents, backlogs, and silent data loss turn a building integration into a long-running engineering problem. I inherited a building controls integration problem 18 months ago. Three office floors. 217 BACnet sensors covering temperature, occupancy, and HVAC actuators. The data was trapped inside the building automation network while the business wanted analytics, reporting, and compliance visibility in the data platform. The obvious answer looked easy enough: deploy bacnet2mqtt, bridge BACnet into MQTT, and push the stream into the lakehouse stack. The repository made it sound like a w...

Get Apache Flume 1.3.x running on Windows

Since we found an increasing interest in the flume community to get Apache Flume running on Windows systems again, I spent some time to figure out how we can reach that. Finally, the good news - Apache Flume runs on Windows. You need some tweaks to get them running. Prerequisites Build system: maven 3x, git, jdk1.6.x, WinRAR (or similar program) Apache Flume agent: jdk1.6.x, WinRAR (or similar program), Ultraedit++ or similar texteditor Tweak the Windows build box 1. Download and install JDK 1.6x from Oracle 2. Set the environment variables    => Start - type " env " into the search box, select " E dit system environment variables ", click Environment Variables, Select " New " from the " Systems variables " box, type " JAVA_HOME " into " variable name " and the path to your JDK installation into "Variable value" (Example:  C:\Program Files (x86)\Java\jdk1.6.0_33 ) 3. Download maven from Apache 4. Set...