Hadoop Cluster Configuration Settings

This chapter contains the following sections:

Creating a Hadoop Cluster Configuration Parameters Template

You can modify the Hadoop Cluster Configuration Parameters Template only from the Hadoop Cluster Configuration Parameters Template tab on the menu bar from Solutions > Big Data > Settings before triggering a Hadoop cluster.


Note


Select the Hadoop cluster configuration parameters template to edit, clone, or delete.



    Step 1   On the menu bar, choose Solutions > Big Data > Settings.
    Step 2   Click the Hadoop Config Parameters tab.
    Step 3   Click Add.
    Step 4   In the Hadoop Config Parameters page of the Create Hadoop Cluster Configuration Parameters Template wizard, complete the following fields:

    Name

    Description

    Template Name field

    A unique name for the Hadoop cluster configuration parameter template.

    Template Description field

    The description for the Hadoop cluster configuration parameter template.

    Hadoop Distribution drop down list

    Choose the Hadoop distribution.

    Hadoop Distribution Version drop down list

    Choose the Hadoop distribution version.

    Step 5   Click Next.
    Step 6   In the Hadoop Config Parameters - HDFS Service page of the Create Hadoop Cluster Configuration Parameters Template wizard, configure the following fields:

    Name

    Description

    HDFS Service

    Hadoop cluster HDFS service parameter name, value, and the minimum supported Hadoop distribution.

    Data Node (role)

    Displays the parameter names, values, and the minimum supported Hadoop distribution version.

    Name Node (role)

    Displays the parameter names, values, and the minimum supported Hadoop distribution version.

    Secondary Node (role)

    Displays the parameter names, values, and the minimum supported Hadoop distribution version.

    Step 7   In the Hadoop Config Parameters - YARN Service page of the Create Hadoop Cluster Configuration Parameters Template wizard, configure the parameters.
    Step 8   In the Hadoop Config Parameters - HBase Service page of the Create Hadoop Cluster Configuration Parameters Template wizard, configure the parameters.
    Step 9   In the Hadoop Config Parameters - MapReduce Service page of the Create Hadoop Cluster Configuration Parameters Template wizard, configure the parameters.
    Step 10   In the Hadoop Config Parameters - Miscellaneous Parameters page of the Create Hadoop Cluster Configuration Parameters Template wizard, configure the (ServiceLevel and RoleLevel) parameters.
    Step 11   Click Submit.

    Updating Hadoop Cluster Configuration Parameters Template - Post Hadoop Cluster Creation


      Step 1   On the menu bar, choose Solutions > Big Data > Accounts.
      Step 2   Click the Big Data Accounts tab and choose an existing Big Data Account.
      Step 3   Click Configure Cluster.
      Step 4   In the Hadoop Config Parameters page of the Update Hadoop Cluster Configuration Parameters Template wizard, complete the following fields:

      Name

      Description

      Hadoop Distribution drop down list

      Choose the Hadoop distribution.

      Hadoop Distribution Version

      Displays the selected Hadoop distribution version.

      Step 5   Click Next.
      Step 6   In the Hadoop Config Parameters - HDFS Service page of the Update Hadoop Cluster Configuration Parameters Template wizard, update the following fields:

      Name

      Description

      HDFS Service

      Hadoop cluster HDFS service parameter name, value, and the minimum supported Hadoop distribution version, if any.

      Data Node (role)

      Displays the parameter names, values, and the minimum supported Hadoop distribution version that can be edited.

      Name Node (role)

      Displays the parameter names, values, and the minimum supported Hadoop distribution version that can be edited.

      Secondary Node (role)

      Displays the parameter names, values, and the minimum supported Hadoop distribution version that can be edited.

      Step 7   In the Hadoop Config Parameters - YARN Service page of the Update Hadoop Cluster Configuration Parameters Template wizard, update the parameters.
      Step 8   In the Hadoop Config Parameters - HBase Service page of the Update Hadoop Cluster Configuration Parameters Template wizard, update the parameters.
      Step 9   In the Hadoop Config Parameters - MapReduce Service page of the Update Hadoop Cluster Configuration Parameters Template wizard, update the parameters.
      Step 10   In the Hadoop Config Parameters - Miscellaneous Parameters page of the Update Hadoop Cluster Configuration Parameters Template wizard, update the (ServiceLevel and RoleLevel) parameters.
      Step 11   Click Submit.

      Quality of Service System Class

      Quality of Service

      For more information on Quality of Service and System Classes, see QoS System Classes.

      Cisco Unified Computing System provides the following methods to implement quality of service (QoS):
      • System classes that specify the global configuration for certain types of traffic across the entire system.

      • QoS policies that assign system classes for individual vNICs.

      • Flow control policies that determine how uplink Ethernet ports handle pause frames.

      System Classes

      Cisco UCS uses Data Center Ethernet (DCE) to handle all traffic inside a Cisco UCS domain. This industry standard enhancement to Ethernet divides the bandwidth of the Ethernet pipe into eight virtual lanes. Two virtual lanes are reserved for internal system and management traffic. You can configure quality of service (QoS) for the other six virtual lanes. System classes determine how the DCE bandwidth in these six virtual lanes is allocated across the entire Cisco UCS domain.

      Each system class reserves a specific segment of the bandwidth for a specific type of traffic, which provides a level of traffic management, even in an oversubscribed system. For example, you can configure the Fibre Channel Priority system class to determine the percentage of DCE bandwidth allocated to FCoE traffic.

      The following table describes the system classes that you can configure

      System Class

      Description

      Best Effort

      A system class that sets the quality of service for the lane reserved for basic Ethernet traffic. Some properties of this system class are preset and cannot be modified.

      For example, this class has a drop policy that allows it to drop data packets if required. You cannot disable this system class.

      • Platinum

      • Gold

      • Silver

      • Bronze

      A configurable set of system classes that you can include in the QoS policy for a service profile. Each system class manages one lane of traffic. All properties of these system classes are available for you to assign custom settings and policies.

      Fibre Channel

      A system class that sets the quality of service for the lane reserved for Fibre Channel over Ethernet traffic. Some properties of this system class are preset and cannot be modified.

      For example, this class has a no-drop policy that ensures it never drops data packets. You cannot disable this system class.

      Note   

      FCoE traffic has a reserved QoS system class that should not be used by any other type of traffic. If any other type of traffic has a CoS value that is used by FCoE, the value is remarked to 0

      Editing QoS System Class

      For more information on Quality of Service and System Classes, see QoS System Classes.


        Step 1   On the menu bar, choose Solutions > Big Data > Settings.
        Step 2   Click the QoS System Class tab.
        Step 3   Choose the QoS System Class (by Priority) that you want to edit and click Edit.
        • Best Effort

        • Platinum

        • Gold

        • Silver

        • Bronze

        Step 4   In the Modify QoS System Class dialog box, complete the following fields:

        Name

        Description

        Enabled check box

        If checked, the associated QoS class is configured on the fabric interconnect and can be assigned to a QoS policy.

        If unchecked, the class is not configured on the fabric interconnect and any QoS policies associated with this class default to Best Effort or, if a system class is configured with a CoS of 0, to the Cos 0 system class.

        This field is always checked for Best Effort and Fibre Channel.

        CoS drop-down list

        The class of service. You can enter an integer value between 0 and 6, with 0 being the lowest priority and 6 being the highest priority. We recommend that you do not set the value to 0, unless you want that system class to be the default system class for traffic if the QoS policy is deleted or the assigned system class is disabled.

        This field is set to 7 for internal traffic and to any for Best Effort. Both of these values are reserved and cannot be assigned to any other priority.

        Packet Drop check box

        This field is always unchecked for the Fibre Channel class, which never allows dropped packets, and always checked for Best Effort, which always allows dropped packets.

        If checked, packet drop is allowed for this class. If unchecked, packets cannot be dropped during transmission.

        Weight drop-down list

        This can be one of the following:

        • An integer between 1 and 10. If you enter an integer, Cisco UCS determines the percentage of network bandwidth assigned to the priority level as described in the Weight (%) field.

        • best-effort.

        • none.

        Muticast Optimized check box

        If checked, the class is optimized to send packets to multiple destinations simultaneously. This option is not applicable to the Fibre Channel.

        MTU drop-down list

        The maximum transmission unit for the channel. This can be one of the following:
        • An integer between 1500 and 9216. This value corresponds to the maximum packet size.

        • fc—A predefined packet size of 2240.

        • normal—A predefined packet size of 1500.

        • Specify Manually—A packet size between 1500 to 9216.

        This field is always set to fc for Fibre Channel.

        Step 5   Click Submit.

        Pre Cluster Performance Testing Settings

        You can analyze memory, network, and disk metrics and a default Big Data Metrics Report provides the statistics collected for each host before creating any Hadoop cluster.


          Step 1   On the menu bar, choose Solutions > Big Data > Settings.
          Step 2   Click the Management tab.
          Step 3   In the Pre Cluster Performance Tests section, check the check boxes for the following:
          • Memory Test

          • Network Test

          • Disk Test

          Note   

          By default, the check boxes to run the memory, network, and the disk tests are unchecked. If you enable the pre-cluster disk test that may impact significantly Hadoop cluster creation.

          Step 4   Click Submit.

          Approving Hadoop Cluster Deployment Workflows

          Before You Begin
          Choose Administration > Users and Groups > Users and add users with the following user roles:
          • Network Admin (system default user role)

          • Computing Admin (system default user role)

          • Hadoop User


            Step 1   On the menu bar, choose Solutions > Big Data > Settings.
            Step 2   Click the Management tab.
            Step 3   Check the Require OS User Approval check box.
            1. From the User ID table, select the Login Name of the user against the Network Admin user role.
            2. Enter the Number of Approval Request Reminders.
              Note   

              Set the number of approval reminders to Zero if the reminder e-mail has to be sent at a specified interval till the Network Admin approves or rejects the approval request.

            3. Enter the Reminder Interval(s) in hours.
            Note   

            Check the Approval required from all the users check box, if you want all the users to approve or reject the approval request.

            Step 4   Check the Require Compute User Approval check box.
            1. From the User ID table, select the Login Name of the user against the Computing Admin user role.
            2. Enter the Number of Approval Request Reminders.
              Note   

              Set the number of approval reminders to Zero if the reminder e-mail has to be sent at a specified interval till the Computing Admin approves or rejects the approval request.

            3. Enter the Reminder Interval(s) in hours.
            Note   

            Check the Approval required from all the users check box, if you want all the users to approve or reject the approval request.

            Step 5   Check the Require Hadoop User Approval check box.
            1. From the User ID table, select the Login Name of the user against the Hadoop User user role.
            2. Enter the Number of Approval Request Reminders.
              Note   

              Set the number of approval reminders to Zero if the reminder e-mail has to be sent at a specified interval till the Hadoop User approves or rejects the approval request.

            3. Enter the Reminder Interval(s) in hours.
            Note   

            Check the Approval required from all the users check box, if you want all the users to approve or reject the approval request.

            Step 6   Click Submit.

            What to Do Next

            Check if users of Network Admin, Computing Admin, and Hadoop User roles have approved the request before deploying any Hadoop cluster.

            Uploading Required OS and Hadoop Software to Cisco UCS Director Baremetal Agent

            You can upload (add) required RHEL 6.x ISO files, Hadoop software and common software that are required for Hadoop distributions to Cisco UCS Director Baremetal Agent. While uploading the required files from your local or any remote system, the files are first uploaded to Cisco UCS Director, and then moved to the target Cisco UCS Director Baremetal Agent once you click the Submit button in the Create Software Catalogs dialog box.

            Supported file formats:
            • Linux OS - rhel-x.x.iso

            • Hadoop software - MapR-x.y.z.zip (.gz or .tgz or .tar)

            • Common software - bd-sw-rep.zip (.gz or .tgz or .tar)

            The Software Catalogs page displays Hadoop distributions and the required software for those Hadoop distributions in the Cisco UCS Director Baremetal Agent.


            Note


            If the required software column is empty for a Hadoop distribution, then it means that Cisco UCS Director Baremetal Agent contains all the files required.



              Step 1   On the menu bar, choose Solutions > Big Data > Settings.
              Step 2   Click the Software Catalogs tab.
              Step 3   Click Add.
              Step 4   Click Upload to upload files from your local system.
              Note   

              You must create a folder for the Hadoop distribution to include all the required files and compress the folder before uploading in the format specified.

              Step 5   Choose the target Cisco UCS Director Baremetal Agent from the Target BMA drop-down list.
              Step 6   Check the Restart BMA Services to restart Cisco UCS Director Baremetal Agent after uploading the required files.
              Note   

              Refresh the Software Catalogs page after 5 to 10 minutes to see new and modified catalogs.

              Linux OS Upload

              Catalog Name field

              Operating System Name (for example, RHEL.6.5)

              Upload Type drop-down list

              Choose one of the following:
              • Desktop file

              • The web server path that is reachable by the Cisco UCS Director Baremetal Agent

              • Mountpoint in Cisco UCS Director Baremetal Agent (For example, /root/iso)

              • Path to ISO in Cisco UCS Director Baremetal Agent(For example, /temp/rhel65/iso)

              Hadoop Software Upload

              Catalog Name field

              Hadoop Distribution (for example, distribution_name-x.y.z)

              Upload Type drop-down list

              Choose one of the following:
              • Desktop file

              • The web server path that is reachable by the Cisco UCS Director Baremetal Agent to upload remote software to Baremetal Agent

              Common Software Upload

              Upload Type drop-down list

              Choose one of the following:
              • Desktop file

              • The web server path that is reachable by the Cisco UCS Director Baremetal Agent to upload remote software to Baremetal Agent

              Step 7   Click Submit.

              What to Do Next

              You can track software uploads here: Administration > Integration. Click the Change Record tab to track the software upload in progress and verify if completed or failed or timeout.

              Cloudera, MapR, and Hortonworks RPMs on Cisco UCS Director Express for Big Data Baremetal Agent

              Common Packages for Cloudera, MapR, and Hortonworks


              Note


              For any Hadoop software that is not available, you have to update the /opt/cnsaroot/bigdata_templates/common_templates/HadoopDistributionRPM.txt with the available online repository of the vendor.



              Note


              We recommend to verify the supported versions from the Hadoop Vendor Support Documentation.


              Download the following common packages to /opt/cnsaroot/bd-sw-rep/:

              Common Packages for Cloudera

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.X.X:

              • ClouderaEnterpriseLicense.lic— Get the license keys from Cloudera

              • userrpmlist.txt—For additional packages list

              • catalog.properties—Provides the label name for the Cloudera version (x represents the Cloudera version on the Cisco UCS Director Express for Big Data Baremetal Agent)

              • mysql-connector-java-5.1.26.tar.gz from http:/​/​cdn.mysql.com/​archives/​mysql-connector-java-5.1

              Cloudera 5.0.1 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.0.1:

              Cloudera 5.0.6 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.0.6:

              Cloudera 5.2.0 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.2.0:

              Cloudera 5.2.1 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.2.1:

              Cloudera 5.3.0 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.3.0:

              Cloudera 5.4.1 Packages and Parcels

              Download the following packages to /opt/cnsaroot/bd-sw-rep/cloudera-5.4.1:

              Common Packages for MapR 3.1.1, 4.0.1, and 4.0.2

              Download the following common packages to /opt/cnsaroot/bd-sw-rep/MapR-3.1.1 and MapR-4.0.x directories:

              Common Packages for MapR 4.1.0 and 5.0.0

              Download the following common packages to /opt/cnsaroot/bd-sw-rep/MapR-4.1.0 and MapR-5.0.0 directories:

              MapR 3.1.1 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/MapR-3.1.1

              MapR 4.0.1 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/MapR-4.0.1

              MapR 4.0.2 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/MapR-4.0.2

              MapR 4.1.0 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/MapR-4.1.0

              MapR 5.0.0 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/MapR-5.0.0

              Common Package for Hortonworks

              Download the following common package to /opt/cnsaroot/bd-sw-rep/Hortonworks-2.X:
              • openssl-1.0.1e-30.el6.x86_64.rpm

              • catalog.properties—Provides the label name for the Hortonworks version (x represents the Hortonworks version on the Cisco UCS Director Express for Big Data Baremetal Agent)

              Hortonworks 2.1 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/Hortonworks-2.1:

              Hortonworks 2.2 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/Hortonworks-2.2:

              Hortonworks 2.3 Packages

              Download the following packages to /opt/cnsaroot/bd-sw-rep/Hortonworks-2.3:

              Cloudera and MapR RPMs for Upgrading Hadoop Cluster Distributions

              Cloudera 5.3.0 Packages and Parcels

              Cloudera 5.4.1 Packages and Parcels

              MapR 4.1.0 Packages

              MapR 5.0.0 Packages

              Configuration Check Rules

              You can validate an existing cluster configuration by running a configuration check. The configuration check process involves comparing the current cluster configuration with configuration check rules and reporting violations.

              Configuration check rules are pre-defined Cisco Validated Design (CVD) parameters for Hadoop clusters. Configuration check rules appear under Solutions > Big Data > Settings. After the configuration check is complete, the violations appear in the Faults tab under Solutions > Big Data > Accounts. You can enable or disable configuration check rules at any time, but you cannot add new rules.

              Configuration Check Rule

              Description

              Parameter

              The pre-defined CVD parameter of the configuration.

              Enabled

              The state of the configuration check rule, either enabled (true) or disabled (false).

              Expected value

              The value expected for a parameter as defined in the Cisco Validated Design (CVD).

              Description

              The description of the parameter of the configuration.

              Distribution

              The Hadoop distribution.

              Minimum Supported Distribution

              The minimum supported version of Hadoop distribution.

              Service

              The Hadoop service.

              Role

              The Hadoop service role.

              Type

              The type of violation, either CVD or Inconsistent.

              Fix Workflow

              The reference to the workflow that can be triggered for fixing violations.

              When the actual cluster configuration values differ from the expected values defined in the configuration check rules, then those configuration values are reported as violations. For example, CVD mandates that the NameNode heap size should be 4 GB. But if the NameNode heap size in the cluster configuration is found to be 1 GB, then this is reported as a CVD violation. Additionally, inconsistent configuration parameters are reported. For example, NameNode heap size on both the primary and secondary nodes must be of the same size. If there is a mismatch in the size, then this parameter is reported as inconsistent.

              Checking Hadoop Cluster Configuration

              To validate the configuration of a cluster, do the following:


                Step 1   On the menu bar, choose Solutioms > Big Data > Accounts.
                Step 2   Click the Big Data Accounts tab.
                Step 3   Choose the account for which you want to run the configuration check and click Check Configuration.
                Step 4   Click Submit.

                A dialog box appears with the information that the configuration check is in progress.

                Step 5   Click OK.

                After the configuration check is complete, the violations appear under the Faults tab for the selected Big Data Account.


                What to Do Next


                Note


                You can track configuration checks here: Administration > Integration. Click the Change Record tab to track the configuration checks in progress and verify if completed or failed.


                Fixing Configuration Violations

                After the configuration check is complete, the configuration violations appear in the Faults tab for the selected BIg Data Account. You can either choose to fix these configuration violations manually on the Big Data Cluster Configuration page or trigger a workflow. To trigger a workflow to fix the violation, first you must create a workflow with the same name as the code specified in the violation.

                To fix a configuration violation through a workflow, do the following:


                  Step 1   On the menu bar, choose Solutions > Big Data > Accounts.
                  Step 2   Click the Faults tab.
                  Step 3   Choose the configuration violation you want to fix and click Trigger Workflow.

                  If a workflow exists with the same name as the code specified in the violation, then the workflow is triggered.

                  Step 4   Enter the required inputs for the workflow and click Submit. A service request ID is generated after you submit the inputs. You can check the status of the service request on the Service Requests page.