We've had an opportunity to man the Centrify booth at the Strata+Hadoop show in New York this past week and we were constantly surprised by the focus on the core capability, but when we engaged passers by, the were unaware of the following facts:
A large majority of Big Data projects never make it into production
The most visible examples of Big Data projects that add value to organizations involve data that may contain Personal Identifiable Information (PII) or financial information.
It does not take a security expert to realize that these projects would be subject to the biggest security requirements.
When challenged, we heard answers like "but [management tool] can be integrated to Active Directory" (on the more informed side of the spectrum) or simply, people were focusing on areas like table-level security.
It became clear to me, that "to secure" a capability has a different meaning depending on who you ask (e.g. Big Data Scientist vs. Linux administrator vs. Security Analyst); this is why at Centrify we always encourage the main stakeholder to bring in their infrastructure and security peers.
people: being able to find capable data scientists
data: availability of the data
cost
dynamics: integration with business processes
need: is it just a fad (business want to jump into the BigData bandwagon without real requirements)
But security will always come up, that's a common denominator across all Big Data. We know this first hand because that's what we do. Upwards of 60 of our customers are using Centrify to align their Big Data deployments with security requirements.
If you're tasked to launch a Big Data project, you have all the opportunity to get a leg-up on this problem and attack security early.
At a basic level you have to look at this in terms of layers.
To simplify, there's two of them: The Identity/Infrastructure layer and the Big Data layer. The concerns and focus are completely different.
OS Layer
I also like to call it the Identity/Infrastructure layer because it plays directly into what Identity and Access Management is set out to do:
a) Use a common directory to identify and authenticate users
b) Enforce the least access principle
c) Enforce the least privilege principle
d) Eliminate the human problem of shared accounts
e) Implement strong controls (e.g. end-to-end session auditing or multi-factor authentication) when needed
f) Be able to attest who has access to a system
g) Provide reporting and tools for attestation
Big Data Layer
I can't even pretend to advise on the Big Data layer, but all I know is this:
If you can't identify - you can't authenticate - if you can't authenticate, you can't authorize; and if you can't authorize, you can't enforce strong access controls. Regardless of Knox, Sentry and any other security initiatives at the Big Data layer, you need robust OS services to optimize those.
A promising future
We were also excited by the bright spots: Cloudera, Hortonworks and MapR are taking security very seriously because they realized that this affects their ability to bring nodes into production.
Centrify is here to help!!! Learn what we mean by "secure" - it's all about the OS Layer (Identity and Access Control)
Overcoming the Hadoop Security Challenges at the IAM layer with Centrify
In the almost two years of Centrifying we have discussed Identity Consolidation with AD and Single Sign-on at length. 90% of organizations have Active Directory but sometimes over-complicate things when it comes to identity consolidation and SSO.
I had the chance to speak about this in a seminar and these two videos consolidate many entries that we've covered in this blog over the years.
Direct Integration
Name Service Switch, Pluggable Authentication Modules, GSSAPI, Kerberos and Proxies
OpenSSH SSO over an outgoing external non-transitive AD one-way trust
NSS and PAM using Oracle DB as an example (externally identified user)
The idea is to eliminate complexity and promote reuse by committing to Active Directory, let the Centrify DirectControl agent do the heavy-lifting for Direct Integration and use the SPNEGO plugins when needed.
For the full briefing, including marketing slideware go here.
This 15-minute video also features Centrify Privileged Service
Centrify can help challenges with Hadoop implementations for confidentiality and integration at the OS-level. No need to stand-up independent MIT Kerberos infrastructure, plus the strongest Access Controls to meet or exceed any security or regulatory requirement.
Hadoop implementations present multiple challenges to enterprises at the Operating System layer(*), I like to categorize them in 2 areas:
Confidentiality
Hadoop clusters are unsecure by default. What that means is that there's no service-to-service authentication and that privileged users have access to world-readable information and can elevate to privileged Hadoop accounts.
Multiple clusters are needed because of the development nature of the apps. Typically at least a DEV/QA and PROD environments are needed, depending on the risk profile of the organization, each environment may be in different isolated environments and require different access control rules.
Different types of users need access: From the SysAdmin, to the Hadoop Admin, to the Data Scientist, they all have different access and privileged needs.
The data classification of the business intelligence may require additional controls. What if the cluster crunches Personal, Financial, Health, Energy or Card data? SOx, PCI, HIPAA, NERC or FERC compliance is needed.
Integration
Kerberos: Many organizations balk at the proposition of standing-up a separate MIT Kerberos implementation; and even if AD is an option, test environments may be in a one-way trust.
Different organizations == Different requirements, therefore the devil is in the details:
Process
Technology/Infrastructure
People
Regulations
(*) There are additional security challenges, like how to protect data at the Hadoop layer, for this, your trusted Hadoop vendor (Cloudera, Hortonworks, MapR, etc) have an ecosystem of applications. Centrify can provide identity information to those apps.
Technical Briefing
The following videos provide technical demos on how Centrify can overcome these challenges
Putting it all together
Centrify allows for OS level integration for Linux and UNIX systems that enables:
Centralized Administrations of multiple Hadoop Clusters
Regardless of how complex your AD may be
No schema extensions or software in domain controllers
Using UNIX frameworks
Kerberos just works out of the box
Leverage AD fully: Kerberos, Group Policy, PKI
Centrify enables the implementation of strong access controls to enforce
Least access
Least privilege (RBAC- not password Centric)
Easy attestation and reporting
Separation of Duties
Works on Windows to eliminate the problem of the persistent administrator
For environments with Personal, Financial, Health or Card data
Session transcription
Session replay
Event consolidation
Works on Windows
Hadoop-exclusive features:
adkeytab for advanced keytab/service account provisioning
Kerberos infinite ticket renewal parameters and GPOs
LDAP Proxy to assist apps like Sentry, Hue and Knox
Partner with Cloudera, MapR and Hortonworks
Centrify + AD + Hadoop = faster, more secure and regulation-aligned big data projects.
In the meantime Centrify has been gearing-up for the release of Centrify Suite 2015, part of what's coming is improvements on all popular Hadoop implementations with Cloudera, Hortonworks and MapR. As a preview, David has released a few of companion whitepapers:
This playlist also explains the challenges at the OS Level:
*********************************************************************************
As of February 2015 Centrify is gearing-up to release Centrify Suite 2015, with it, enhancements specific to Hadoop deployments are being released. To complement these efforts, integration guides for Cloudera, Hortonworks and MapR are being released, here are the first two:
These are more robust and well-researched papers. The article below from August 2014 will be left there for history purposes, but the post and the accompanying video outdated.
*********************************************************************************
Background
Big Data is one of the fastest spreading IT trends in the enterprise today, it's also a big reason why IT infrastructures have to provide elastic and secure infrastructures. Apache Hadoop (with value-added by Cloudera, Hortonworks, MapR, IBM and others) is gaining a lot of traction in enterprises that are looking to adopt Big Data.
What does this mean to the IT Infrastructure Manager?
More Linux servers
Need for more effective management
More systems that need to be aligned with security practices
Depending data classification, a mature access management model that allows the enforcement of the principles of least access, least privilege, separation of duties, quick attestation and optionally session capture and replay is required.
Customers of Centrify know that all the bullets above are not an issue with properly Centrified (*) systems.
(*) A mature Centrify deployment has no systems in Express mode, has discontinued the use of Classic zones and archaic provisioning methods like ZoneGen, and finally is using RBAC via Centrify-enhanced sudo. Recommendations:
If you are not familiar with terms like Kerberos authentication, keytabs, servicePrincipalName (SPN), userPrincipalName (UPN) and some of the differences between MIT Kerberos and Microsoft Kerberos, I recommend that you familiarize with those topics.
From a Hadoop perspective, leverage Hortonworks/Cloudera/MapR/IBM's professional services organizations. Hadoop is a broad topic and in my observation successful implementations have one thing in common: they get expert help. You are busy enough as it is just keeping-up with the day-by-day.
Keep in mind, the reason why we introducing Active Directory as an alternative to a stand-alone MIT Kerberos is for a simple reason: to eliminate duplication of capabilities. If you go the MIT Kerberos route, this means that you have to maintain two different technology environments + process + people implications. This just doesn't make sense.
If you feel comfortable with the topics above, you at least need to understand what the adkeytab utility is. This post covers adkeytab.
Centrify & AD = Faster Hadoop Implementation
If you've been following this blog, you probably realize the focus on Access Controls and to leverage your existing infrastructure (AD); this extends to Hadoop deployments as well. Because:
You already would have solved the Access and Privilege governance model at the OS level.
You don't need to add more complexity to your environment to enable Hadoop Security.
If you don't have a mature access and privileged model like local accounts or shared accounts, you don't want to continue to propagate the problem.
Hadoop leverages MIT Kerberos to enable security by Kerberizing its services. With Kerberos no passwords go through the wire and there are compensating controls for confidentiality, integrity and other threats.
What is the impact?
Hadoop providers will ask you to stand-up an MIT Kerberos environment to support your BigData deployment; a very simple ask, but with a large impact to the enterprise. Things to take into account:
Kerberos infrastructure - additional services that need to be highly available.
Additional process - add/moves/changes of principals as the nodes expand/contract/change
Potentially a Kerberos trust: If you want to leverage AD, you may need into the business of a Kerberos trust into AD. At least one provider has a recipe for this.
Push-back: Believe it or not, the idea of an additional authentication infrastructure is not going to be welcome by some IT groups.
This is where Centrify can help. Centrify provides the Kerberos environment and tools, services and utilities to extend your access model and accelerate the Hadoop deployment. With that in mind, let's apply the Plan-Do-Check-Adjust model.
Note: Keep in mind, Hadoop follows the MIT kerberos implementation to the letter. This means that there's no concept of multiple service principal names tied to an account (Unlike AD), therefore as of today, that implies that some services will require one account per service per node. I repeat now in AD terms: One AD Account/Keytab, Per Service, Per Node. This is where utilities like adkeytab can help in the automation process. Number of Accounts (Keytabs) = Shared Service Accounts + (Nodes * Unique Service Accounts) E.g. If you have a 50-node cluster that has 2 services that can share the same account/keytab and 5 services that require unique accounts/keytabs, this means that you will have 2 + 50*5 = 252 AD accounts/unique keytabs.
Believe it or not, the easy part is the technology, the hard part is the process and the automation required.
Planning to Secure your Hadoop Services with Centrify
Infrastructure vs. User Apps:
Kerberizing the servies like HDFS, MapReduce, Hive, Zookeeper, etc is just one piece of the puzzle, the user-level apps may have different access models. Many use OS users and groups and Centrify will have you covered, but this is a different design session.
What is the access model at the OS level? (may imply a different zone or set of computer groups)
What are the privileges (commands) for a Hadoop operator? (remember, with Centrify there's no need to share privileged account credentials) - continue to enforce the least privilege model with Centrify RBAC
Just like the Centrify agent, deploying Name Server Cache Daemon can improve performance.
Have a rock-solid naming convention for:
AD account names: remember the 20 character limit. Make sure that you can identify these accounts correctly in AD. Use a dedicated OU if you have a large cluster. Something like Environment-server-service general-to-specific can work. E.g. "mktqa-had01-spnego"
Keytab location and naming:
Keytab protocol:
Secure based on requirements at rest.
Always transport securely
Keep keytabs where they are needed ONLY.
Will the user accounts in AD have non-expiring randomized passwords or will the passwords be randomized after a period of time?
Does the account used for AD Service account/keytab generation have the proper rights in AD to create those users/principals?
What will be the automation strategy? Remember, using cleartext passwords in scripts or other facilities is a NO-NO, but adkeytab is kerberized; therefore you may need to provision a service account with the proper rights in AD and a keytab to be able to launch adkeytab in the automation script.
How are you implementing Hadoop? Is your team being trained? Succesful implementations (just like with Centrify) make use of consulting services; so leverage your Hortonworks, Cloudera or IBM experts. After all, how many Hadoop (or Centrify) deployments have you implemented? How much pure Kerberos knowledge do you have?
Implementation
In this example, we'll use Hortonworks Hadoop Ambari 1.6.x to enable Kerberos in a cluster. The domain is corp.contoso.com, the OU for service accounts is UNIX\Hadoop and the naming convention is servername-servicename.
Set up your system and join it to AD. Make sure you don't register an SPN for http because it will be used by Hadoop. This can be done: - Prior to the join: by editing the /etc/centrifydc/centrifydc.conf and modifying the adclient.krb5.service.principals directive and removing http. E.g. - After the join: Remove the SPN with ADUC or with adkeytab with the -x (delspn) option. adkeytab --delspn --principal http/hadoop1.corp.contoso.com -V Remember that you can check the service principals with the "adinfo -C" command.
Set up your cluster and make sure your services are running.
Go to Administration-Security and Click Enable Kerberos.
Click Next in the Getting started page.
In the Configure Pages page type the AD domain name in all caps in the Realm Name and in the Kerberos tool path, type the Centrify-enhanced Kerberos tools path (/etc/centrifydc/kerberos/bin).
In the Create Principals and Keytabs page, you'll be shown a table with all the principals per host based on your configuration. Ambari provides a CSV file that can be consumed by a keytab-generating script. In the case of Centrify, you can leverage adkeytab to perform these operations. Here are a few examples of what needs to be done. Note: This example assumes that the person running adkeytab has the rights to create the user principals in Active Directory or in the respective container.
Shared Keytab Example (Ambari) Let's analyze this. a) We will be creating a new account (keep in mind that we can adopt an existing principal) (option --new) b) We will be specifying a UPN (because this keytab will be used to get a Kerberos TGT) (option --upn; e.g. ambari-qa@CORP.CONTOSO.COM) c) We will create a keytab file (option --keytab /path/to/file) d) The AD account will be located in a particular OU (option --container <x.500 notation>). E.g. --container "ou=service-accounts,ou=Unix" e) Verbose output is always recommended (-V) f) The final parameter is the dn (distinguished name) of the account in AD. (e.g. hadoop-ambari-qa) Sample command: adkeytab --new --upn ambari-qa@CORP.CONTOSO.COM --keytab /etc/security/keytabs/smokeuser.headless.keytab --container "ou=hadoop,ou=unix" -V hadoop-ambari-qa Then, the keytab has to be permissioned accordingly: chown ambari-qa:hadoop /etc/security/keytabs/smokeuser.headless.keytab chmod 440 /etc/security/keytabs/smokeuser.headless.keytab
Individual Host Keytab Example (HTTP for hadoop2) Analysis a) We will be creating a new account (keep in mind that we can adopt an existing principal) (option --new) b) We will be specifying a UPN and an SPN (because this keytab will be used to get a Kerberos TGT and a TGS) (options --upn & --principal; e.g. HTTP/hadoop2.corp.contoso.com@CORP.CONTOSO.COM) c) We will create a keytab file (option --keytab /path/to/file) d) The AD account will be located in a particular OU (option --container <x.500 notation>). E.g. --container "ou=service-accounts,ou=Unix" e) Verbose output is always recommended (-V) f) The final parameter is the dn (distinguished name) of the account in AD. (e.g. hadoop-ambari-qa) Sample command: adkeytab --new --upn HTTP/hadoop2.corp.contoso.com@CORP.CONTOSO.COM --principal HTTP/hadoop2.corp.contoso.com@CORP.CONTOSO.COM --keytab /etc/security/keytabs/spnego.service.keytab --container "ou=hadoop,ou=unix" -V hadoop2-http Then, the keytab has to be permissioned accordingly: chown root:hadoop /etc/security/keytabs/spnego.service.keytab chmod 440 /etc/security/keytabs/spnego.service.keytab
Verify that the principals have been created in AD and that the keytabs are in the selected directory with the proper ownership and permissions. In addition, you can use the /usr/share centrifydc/kerberos/bin/kinit -kt command to test the principals for TGT or TGS requests.
Go back to Ambari's security wizard and Apply the changes.
Monitor the services, depending on the performance of your cluster, you may have to start some services manually.
Checking the Implementation
Aside from checking the cluster status, if you're using scripts to automate the add/move/changes of nodes, you need to make sure that those scripts are rock solid.
Adjusting the Implementation
There are many improvements to be gained, and this depends on your security needs and the variety of environments. The most notable is the user-facing apps. Remember that as new environments are spun up and versions of Hadoop change, this process has to be revisited.