Skip to main content

Posts

Manage rights in OpenStack

Openstack lacks on sophisticated rights management, the most users figure. But that's not the case, role management in Openstack is available. First users and groups needs to be added to projects, this can be done per CLI or GUI [1]. Lets say, a group called devops shall have the full control about OpenStack, but others not in that group can have dedicated operation access like create snapshot, stop / start / restart an instance or looking at the floating IP pool. Users, Groups and Policies OpenStack handles the rights in a policy file in /etc/nova/policy.json , using roles definitions per group assigned to all tasks OpenStack provides. It looks like: { "context_is_admin": "role:admin", "admin_or_owner": "is_admin:True or project_id:%(project_id)s", "default": "rule:admin_or_owner", ... } and describes the default - a member of a project is the admin of that project. To add additional rules, they have to be defined h...

Handling Corrupted Kafka Messages and Offset Recovery in Distributed Systems

This article explains how corrupted Kafka messages occurred in early Kafka versions, how offsets were stored in Zookeeper and how to manually recover a stuck consumer. It documents the race condition described in KAFKA 2477, shows how to inspect offsets using Kafka tools or Zookeeper and describes code based and operational strategies for skipping bad messages in older distributed log systems. Handling Corrupted Kafka Messages and Offset Recovery in Distributed Systems In older Kafka deployments, especially versions before 0.9, it was possible for a message in a topic to become unreadable due to corruption. This happened most often when third party frameworks interacted with Kafka internals or when the consumer logic encountered a rare race condition. One such condition was documented in KAFKA 2477, where a lock on Log.read was missing at the consumer level while Log.write remained protected. Under specific timing, this resulted in a corrupted message being written to ...

Building Modern Hyper-Converged Data Platforms with OpenStack and Hadoop

This article explains how hyper-converged data platforms built with OpenStack provide a flexible and scalable foundation for Hadoop and streaming workloads. It covers the differences between static and on-demand Hadoop clusters, the role of HDFS on block storage, how network design and storage layout affect performance, and why in-memory layers like Alluxio can accelerate analytical and IoT workloads. The piece also outlines best-practice architectures for compute, storage, and networking in modern private and hybrid data platforms. Hyper-converged infrastructures have become a mainstream choice for enterprise data platforms. Back in 2016, more than half of surveyed companies were already adopting HCI. Today, the trend has continued, especially as organizations need elastic compute and storage for Hadoop, Spark, and new streaming workloads. Hyper-Converged Data Platforms for Hadoop and Streaming Workloads Hadoop and modern analytical stacks benefit from flexible resource ...

SolR, NiFi, Twitter and CDH 5.7

Since the most interesting Apache NiFi parts are coming from ASF [1] or Hortonworks [2], I thought to use CDH 5.7 and do the same, just to be curious. Here's now my 30 minutes playground, currently running in Googles Compute. On one of my playground nodes I installed Apache NiFi per mkdir /software && cd /software &&  wget http://mirror.23media.de/apache/nifi/0.6.1/nifi-0.6.1-bin.tar.gz   && tar xvfz nifi-0.6.1-bin.tar.gz Then I've set only nifi.sensitive.props.key property in conf/nifi.properties to an easy to remember secret. The next bash /software/nifi-0.6.1/bin/nifi.sh install installs Apache NiFi as an service. After log in into Apache NiFi's WebUI, download and add the template [3] to Apache NiFi, move the template icon to the drawer, open it and edit the twitter credentials to fit your developer account. To use an  schema-less SolR index (or Cloudera Search in CDH) I copied some example files over into a local directory: cp -r ...

Hive on Tez: Why It Was Faster and Why Manual CDH Integration Is Now Legacy

Apache Tez was introduced as a faster, DAG-based execution engine for Hive and other Hadoop workloads, delivering 30–50% speedups over classic MapReduce in many ETL pipelines. This article explains what Tez brought to Hive, how it fit into CDH-era deployments, and why the old practice of hand-compiling Tez against CDH 5.x is now a legacy pattern rather than a recommended approach. Apache Tez was designed as a low-latency, DAG-based execution engine for Hadoop. It replaced many of the heavyweight MapReduce patterns used by early Hive deployments with more efficient execution graphs, reusing containers and avoiding unnecessary materialization steps. In practical terms, switching Hive from MapReduce to Tez often yielded 30–50% faster ETL and reporting jobs on the same hardware, especially for complex multi-stage queries. What Tez Brought to Hive DAG execution : instead of chaining MapReduce jobs, Tez represents the query plan as a directed acyclic graph of tasks. Contai...

Building RPMs with Maven for Reliable Software Deployment

This article explains how to package applications as revisionable RPM artifacts using Maven, a technique that remains valuable for DevOps and platform engineering teams who operate in controlled, reproducible, or air-gapped environments. It covers prerequisites, rpm-maven-plugin configuration, directory mappings, permissions, and how to integrate RPM creation into CI pipelines to deliver consistent deployment units for Java services. Modern DevOps Packaging: Building RPMs with Maven for Reliable Software Deployment In modern DevOps and platform engineering, one of the most underrated tools is still the RPM. Even with the rise of containers, many organizations rely on RPM-based delivery to manage internal services, JVM applications, and deployment flows in secure or air-gapped environments. A revisionable, reproducible, OS-native package is often the cleanest way to promote artifacts through development, staging, and production. Back in 2015, the motivation was simple: ...

How Spark Integrates with Hive Today (and Why Early CDH Versions Required Manual Setup)

Modern Spark integrates with Hive through the SparkSession catalog, allowing unified access to Hive tables without manual classpath or configuration hacks. Earlier CDH 5.x deployments required copying hive-site.xml and adjusting classpaths because Hive on Spark was not fully supported. This updated guide explains the current approach and provides historical context for engineers maintaining legacy clusters. In modern Hadoop and Spark deployments, Spark connects to Hive through the SparkSession catalog. Hive metastore integration is stable, supported and no longer requires manual configuration steps such as copying hive-site.xml or modifying executor classpaths. Using Hive from Spark Today Create a SparkSession with Hive support enabled: val spark = SparkSession.builder() .appName("SparkHive") .enableHiveSupport() .getOrCreate() Once enabled, Spark can query Hive tables directly: spark.sql("SELECT COUNT(*) FROM sample_07").show() Spark hand...

Setting Up MIT Kerberos ↔ Active Directory Cross-Realm Trust for Secure Hadoop Clusters

This post explains how to configure a secure cross-realm Kerberos trust between a MIT KDC and Active Directory for Hadoop environments. It covers modern Kerberos settings, realm definitions, encryption choices, KDC configuration, AD trust creation, and Hadoop’s auth_to_local mapping rules. A final section preserves legacy compatibility for older Windows Server versions, ensuring the article can be used across mixed enterprise environments. Integrating Hadoop with enterprise identity systems often requires establishing a cross-realm Kerberos trust between a local MIT KDC and an Active Directory (AD) domain. This setup allows Hadoop services to authenticate users from AD while maintaining a separate Hadoop-managed realm. We walk through a full MIT Kerberos ↔ AD trust configuration using a modern setup, while preserving legacy notes for older Windows environments still found in long-lived clusters. Example Realms Replace these with your actual realms and hosts: ALO.LOCA...