Sunday, April 29, 2018

Domain where Hadoop can be used: TELECOM


Analyze call detail records (CDRs)


Telcos perform forensics on dropped calls and poor sound quality, but call detail records flow in at a rate of millions per second. This high volume makes pattern recognition and root cause analysis difficult, and often those need to happen in real-time, with a customer waiting for answers. Delay causes attrition and harms servicing margins.
Hortonworks DataFlow (HDF™) can ingest millions of CDRs per second into Hortonworks Data Platform, where Apache™ Storm or Apache Spark™ can process them in real-time to identify troubling patterns. HDP facilitates long-term data retention for root cause analysis, even years after the first issue. This CDR analysis can be used to continuously improve call quality, customer satisfaction and servicing margins.

Service equipment proactively


Transmission towers and their related connections form the spinal chord of a telecommunications network. Failure of a transmission tower can cause service degradation. Replacement of equipment is usually more expensive than repair. There exists an optimal schedule for maintenance: not too early, nor too late.
HDP stores unstructured, streaming, sensor data from the network. Telcos can derive optimal maintenance schedules by comparing real-time information with historical data. Machine learning algorithms can reduce both maintenance costs and service disruptions by fixing equipment before it breaks.

Rationalize infrastructure investments


Telecom marketing and capacity planning are correlated. Consumption of bandwidth and services can be out of sync with plans for new towers and transmission lines. This mismatch between infrastructure investments and the actual return on investment puts revenue at risk.
Network log data helps telcos understand service consumption in a particular state, county or neighborhood. They can then analyze network loads more intelligently (with data stretching over longer periods of time) and plan infrastructure investments with more precision and confidence.

Recommend next product to buy (NPTB)


Telecom product portfolios are complex. Many cross-sell opportunities exist for the installed customer base, and sales associates use in-person or phone conversations to guess about NPTB recommendations, with little data to support their recommendations.
HDP gives a telco the ability to make confident NPTB recommendations, based on data from all of its customers. Confident NPTB recommendations empower sales associates (or self service) and improve customer interactions. An Apache Hadoop® data lake reduces sales friction and creates NPTB competitive advantage similar to Amazon’s advantage in eCommerce.

Allocate bandwidth in real time


Certain applications hog bandwidth and can reduce service quality for others accessing the network. Network administrators cannot foresee the launch of new hyper-popular apps that cause spikes in bandwidth consumption and then slow performance. Operators must respond to bandwidth spikes quickly, to reallocate resources and maintain SLAs.
Streaming data through HDF into HDP for real-time analysis can help network operators visualize spikes in call center data and nimbly throttle bandwidth. Text-based sentiment analysis on call center notes can also help understand how these spikes impact customer experience. This insight helps maintain service quality and customer satisfaction, and also informs strategic planning to build smarter networks.

Develop new products


Mobile devices produce huge amounts of data about how, why, when and where they are used. This data is extremely valuable for product managers, but its volume and variety make it difficult to ingest, store and analyze at scale. Not all data is stored for conversion into business insight. Even the data that is stored may not be retained for its entire useful life.
Apache Hadoop can put rich product-use data in the hands of product managers, which speeds product innovation. It can capture product insight specific to local geographies and customer segments. Immediate big data feedback on product launches allows PMs to rescue failures and maximize blockbusters.

Domain where Hadoop can be used:HEALTHCARE

Access genomic data for new cancer treatments

If we read that a given drug is “40% effective in treating cancer,” another interpretation could be that the drug is 100% effective for patients with a certain genetic profile. However, genomic data is Big Data. The data in a single human genome includes approximately 20,000 genes. Stored in traditional data platforms, this is the equivalent of several hundred gigabytes. Combining each genome with one million variable DNA locations produces the equivalent of about 20 billion rows of data per person.

Researchers at major universities and teaching hospitals are performing big data analytics in genomics with Hortonworks Data Platform as the cost-effective, reliable platform for storing genomic data and combining that with other data on demographics, trial outcomes, and real-time patient responses. They are adopting Hortonworks DataFlow to stream that data into HDP for real-time decisions and long-term cohort analyses. Connected Data Platforms help those doctors learn which drugs and treatments work best for groups of patients across the genetic spectrum.

Monitor patient vitals in real time

In a typical hospital setting, nurses do rounds and manually monitor patient vital signs. They may visit each bed every few hours to measure and record vital signs but the patient’s condition may decline between the time of scheduled visits. This means that caregivers often respond to problems reactively, in situations where arriving earlier may have made a huge difference in the patient’s wellbeing.

New wireless sensors can capture and transmit patient vitals far more frequently than human beings can visit the bedside, and these measurements can stream into a Hadoop cluster. Caregivers can use these signals for real-time alerts to respond more promptly to unexpected changes. HDP uses this data accumulated over time for healthcare predictive analytics, feeding algorithms that proactively help predict the likelihood of an emergency even before it could be detected with a bedside visit.

Reduce cardiac re-admittance rates

Patients with heart disease can be closely monitored while they are in a hospital, but when those patients go home, they may skip their medications or ignore dietary and self-care instructions given by their doctor when they left the hospital.

Congestive heart failure causes fluid retention, which leads to weight gain. In one innovative program at UC Irvine Health, patients could return home with a wireless scale and weigh themselves at regular intervals. Algorithms running in Hortonworks’ healthcare predictive analytics determined unsafe weight gain thresholds and alerted a physician to see the patient proactively, before an emergency re-admittance was necessary.

Machine learning to screen for autism with in-home testing

Autism spectrum disorders affect 1 in 100 children at an annual cost estimated at more than $100 billion. The condition can be detected through behavior at eighteen months, but more than 1 in 4 cases are still undiagnosed at 8 years of age. A small number of clinical testing facilities are oversubscribed, with long wait lists. The most common diagnostic test typically takes 2.5 hours to administer and score.

Dr. Dennis Wall is Director of the Computational Biology Initiative at the Harvard Medical School. In this presentation, he describes a process his team developed for low-cost, mobile screening for autism. It takes less than five minutes and relies on the ability to store large volumes of semi-structured data from brief in-home tests administered and submitted by parents. Wall’s lab also used Facebook to capture user-reported information on autism.

Artificial intelligence running on those huge data sets helps maximize efficiency of diagnosis without loss of accuracy. This approach, in combination with data storage on a Hadoop cluster, can be used for other innovative machine learning diagnostic processes.

Domain where Hadoop can be used: FINANCIAL SERVICES

Screen New Account Applications for Risk of Default


Every day, large retail banks take thousands of applications for new checking and savings accounts. Bankers that accept these applications consult 3rd-party risk scoring services before opening an account. They can (and do) override do-not-open recommendations for applicants with poor banking histories. Many of these high-risk accounts overdraw and charge-off due to mismanagement or fraud, costing banks millions of dollars in losses. Some of this cost is passed on to the customers who responsibly manage their accounts.

Hortonworks Data Platform can store and analyze multiple data streams and help regional bank managers apply predictive analytics to control new financial account risks in their branches. They can match banker decisions with the risk information presented at the time of decision, to control risk by sanctioning individuals, updating policies, and identifying patterns of fraud. Over time, the accumulated data informs algorithms that may detect subtle, high-risk behavior patterns unseen by the bank’s risk analysts.

Monetize Anonymous Banking Data in Secondary Markets


Banks possess massive amounts of operational, transactional and balance data that holds information about macro-economic trends. This information can be valuable for investors and policy-makers outside of the banks, but regulations and internal policies require that these uses strictly protect the anonymity of bank customers.

Retail banks have turned to Hortonworks Data Platform as a common cross-company data lake for data from different LOBs: mortgage, consumer banking, personal credit, wholesale and treasury banking. Both internal managers and consumers in the secondary market derive value from the data. A single point of data management allows the bank to operationalize security and privacy measures such as de-identification, masking, encryption, and user authentication.

Maintain Sub-Second SLAs with a Hadoop “Ticker Plant”


Ticker plants collect and process massive data streams on stock trades, displaying prices for traders and feeding computerized trading systems fast enough to capture opportunities in seconds. Applying predictive analytics to the financial markets is useful for making real-time decisions, and years of historical market data can also be stored for long-term analysis of market trends.

One Hortonworks customer re-architected its ticker plant with HDP as its cornerstone. Before Hadoop, the ticker plant was unable to hold more than ten years of trading data. Now every day gigabytes of data flow in from thousands of server log feeds. This data is queried more than thirty thousand times per second, and Apache HBase enables super-fast queries that meet the client’s SLA targets. All of this, and also a retention horizon extended beyond ten years.

Analyze Trading Logs to Detect Money Laundering


Another Hortonworks customer that provides investment services processes fifteen million transactions and three hundred thousand trades every day. Because of storage limitations, the company used to archive historical trading data, which limited that data’s availability. In the near term, each day’s trading data was not available for risk analysis until after close of business. This created a window of time with unacceptable risk exposure to money laundering or rogue trading.

Now Hortonworks Data Platform supports their AML software and accelerates the firm’s speed-to-analytics and also extends its data retention timeline. A shared data repository across multiple LOBs provides more visibility into all trading activities. The trading risk group accesses this shared data lake to processes more position, execution and balance data. They can do this analysis on data from the current workday, and it is highly available for at least five years—much longer than before.

Domain where Hadoop can be used: ADVERTISING


Mine POS data to identify high-value shoppers


One marketing analytics company specializes in gathering insight at the checkout counter, across many grocers and drug stores. They mine this sales information for basket analysis, price sensitivity, and demand forecasts.
Interactive query with the Stinger Initiative and Apache Hive running on YARN in hadoop helped the company rapidly process terabytes of data to keep pace with a market that changes by the day. Manufacturers, retailers, and ad agencies use the combined analysis to position their brands or improve their retail experiences, particularly for high-value customers.

Target ads to customers in specific cultural or linguistic segments


Hortonworks customer Luminar is the leading big data analytics and modeling provider uniquely focused on delivering actionable advertising insights on U.S. Latino consumers. Luminar wanted to move beyond sample data on Latino consumers living in the United States, towards empirical analysis of actual data on all US Latinos. Rather than store only some transactions from one or two sources, they wanted to acquire and save as many transactions as possible from as many different sources as possible.
Now hadoop interacts easily with other components of Luminar’s data and business intelligence ecosystem: Amazon Cloud, R, Talend and Tableau. The company has increased ingest of transaction data from 300 sources to thousands, up from 2 to 15 terabytes per month. Before, it took Luminar days to ingest and join a new set of raw data, now it takes only hours, even with eight times more data than before. Luminar uses their actionable intelligence to craft marketing strategies for CPG and entertainment companies that want to focus on the US Latino population.

Syndicate videos according to behavior, demographics and channel


A major omni-media company specializes in home improvement and DIY content distributed across television, digital, mobile and publishing channels. One of its divisions is focused on delivering online video ads.
Both content syndicators and publishers want to make sure that video content reaches the right audience. The company analyzes clickstream data stored on hadoop to analyze audiences then feed insight into a recommendation engine that improves advertising click-through.

ETL toy market research data for longer retention and deeper insight


A leading consumer research firm provides consumer intelligence to the toy industry. The market is in a state of flux; there are more new digital options and “real world” forms of children’s play than ever before. The company helps its clients keep pace by delivering weekly point-of-sale (POS) tracking information for competitive insight on toy sales trends. They cover all the major toy retailers for a single view of the marketplace.
The company chose hadoop to offload much of its data from a more expensive platform, with expected savings of more than $1 million annually. The improved economics allow the company to retain data longer and identify long-range, strategic opportunities for growth. This helps its clients in the toy industry partner more closely with retailers.

Optimize online ad placement for retail websites


An advertising Hortonworks customer provides web analytics services to some of the world’s largest retail websites. For their largest customer, clickstream data pours in at the rate of hundreds of megabytes per hour, which adds up to billions of rows per month. The agency analyzes each ad’s placement and determines click-through and conversion rates. When impression files and click files were stored in a relational database, the agency had no way to intelligently connect impressions to clicks–so they had to make too many guesses.
Now HDP® replaces that guess work with empirical science and confident analysis by week, by day or by hour. The agency can also filter by the consumer’s OS, browser, device and geographical location. With Hortonworks Data Platform’s economies of scale, data storage costs are significantly lower than before, and data can be retained for longer. Now the agency and its clients all look forward to looking back on years (not weeks) of clickstream data.
The agency’s retail customers can now tell if consumers are clicking on their website while standing in one of their stores. This provides valuable insight to manage “showrooming” behavior where customers visit a store to touch a product and then drive home to buy it online. Retailers can address showrooming without slashing prices, and data in HDP reveals specific tactics for doing so.

Thursday, March 15, 2018

Apache Ranger

HDFS is core part of any Hadoop deployment and in order to ensure that data is protected in Hadoop platform, security needs to be baked into the HDFS layer. HDFS is protected using Kerberos authentication, and authorization using POSIX style permissions/HDFS ACLs or using Apache Ranger.

Apache Ranger is a centralized security administration solution for Hadoop that enables administrators to create and enforce security policies for HDFS and other Hadoop platform components.

How Ranger policies work for HDFS?

In order to ensure security in HDP environments, we recommend all of our customers to implement Kerberos, Apache Knox and Apache Ranger.

Apache Ranger offers a federated authorization model for HDFS. Ranger plugin for HDFS checks for Ranger policies and if a policy exists, access is granted to the user. If a policy doesn’t exist in Ranger, then Ranger would default to native permissions model in HDFS (POSIX or HDFS ACL). This federated model is applicable for HDFS and Yarn service in Ranger.



For other services such as Hive or HBase, Ranger operates as the sole authorizer which means only Ranger policies are in effect. The option for the fallback model is configured using a property in Ambari → Ranger → HDFS config → Advanced ranger-hdfs-security

The federated authorization model enables customers to safely implement Ranger in an existing cluster without affecting jobs which rely on POSIX permissions. We recommend enabling this option as the default model for all deployments.

Ranger’s user interface makes it easy for administrators to find the permission (Ranger policy or native HDFS) that provides access to the user. Users can simply navigate to Ranger→ Audit and look for the values in the enforcer column of the audit data. If the populated value in Access Enforcer column is “Ranger-acl”, it indicates that a Ranger policy provided access to the user. If the Access Enforcer value is “Hadoop-acl”, then the access was provided by native HDFS ACL or POSIX permission.


BEST PRACTICES FOR HDFS AUTHORIZATION
Having a federated authorization model may create a challenge for security administrators looking to plan a security model for HDFS.

After Apache Ranger and Hadoop have been installed, we recommend administrators to implement the following steps:

Change HDFS umask to 077
Identify directory which can be managed by Ranger policies
Identify directories which need to be managed by HDFS native permissions
Enable Ranger policy to audit all records
Here are the steps again in detail.

1. Change HDFS umask to 077 from 022. This will prevent any new files or folders to be accessed by anyone other than the owner
Administrators can change this property via Ambari:


The umask default value in HDFS is configured to 022, which grants all the users read permissions to all HDFS folders and files. You can check by running the following command in recently installed Hadoop

$ hdfs dfs -ls /apps
Found 3 items
drwxrwxrwx   – falcon hdfs       0 2015-11-30 08:02 /apps/falcon
drwxr-xr-x   – hdfs   hdfs           0 2015-11-30 07:56 /apps/hbase
drwxr-xr-x   – hdfs   hdfs           0 2015-11-30 08:01 /apps/hive

2. How to identify the directories that can be managed by Ranger policies?

We recommend that permission for application data folders (/apps/hive, /apps/Hbase), as well as any custom data folders, be managed through Apache Ranger. The HDFS native permissions for these directories need to be restrictive. This can be done through changing permissions in HDFS using chmod.

Example:

$ hdfs dfs -chmod -R 000 /apps/hive
$ hdfs dfs -chown -R hdfs:hdfs /apps/hive
$ hdfs dfs -ls /apps/hive
Found 1 items
d———   – hdfs hdfs          0 2015-11-30 08:01 /apps/hive/warehouse

Then navigate to Ranger admin and give explicit permission to users as needed. For example:


Administrators should follow the same process for other data folders as well. You can validate  whether your changes are in effect by doing the following:

Connect to HiveServer2 using beeline
Create a table
create table employee( id int, name String, ssn String);
Go to the ranger, and check the HDFS access audit. The enforcer should be ‘ranger-acl’

3. Identify directories which can be managed by HDFS permissions. It is recommended to let HDFS manage the permissions for /tmp and /user folders. These are used by applications and jobs which create user level directories.

Here, you should also set the initial permission for /user folder  to “700”, similar to the example below

 hdfs dfs -ls /user
Found 4 items
drwxrwx—   – ambari-qa hdfs          0 2015-11-30 07:56 /user/ambari-qa
drwxr-xr-x   – hcat      hdfs          0 2015-11-30 08:01 /user/hcat
drwxr-xr-x   – hive      hdfs          0 2015-11-30 08:01 /user/hive
drwxrwxr-x   – oozie     hdfs          0 2015-11-30 08:02 /user/oozie

$ hdfs dfs -chmod -R 700 /user/*
$ hdfs dfs -ls /user
Found 4 items
drwx——   – ambari-qa hdfs          0 2015-11-30 07:56 /user/ambari-qa
drwx——   – hcat      hdfs          0 2015-11-30 08:01 /user/hcat
drwx——   – hive      hdfs          0 2015-11-30 08:01 /user/hive
drwx——   – oozie     hdfs          0 2015-11-30 08:02 /user/oozie

4.  Ensure auditing for all HDFS data.
Auditing in Apache Ranger can be controlled as a policy. When Apache Ranger is installed through Ambari, a default policy is created for all files and directories in HDFS and with auditing option enabled.This policy is also used by Ambari smoke test user “ambari-qa” to verify HDFS service through Ambari. If administrators disable this default policy, they would need to create a similar policy for enabling audit across all files and folders.

Summary:

Securing HDFS files through permissions is a starting point for securing Hadoop. Ranger provides a centralized interface for managing security policies for HDFS. Security administrators are recommended to use a combination of HDFS native permissions and Ranger policies to provide comprehensive coverage for all potential use cases. Using the best practices outlined in this blog, administrators can simplify the access control policies for administrative and user directories, files in HDFS.

Thursday, March 1, 2018

Hadoop Admin Interview Question Answer -3

Q 1. In Hadoop ecosystem, we have HDFS, Zookeeper, Yarn/ Mapreduce2, Hive, spark, oozie. What is the sequence of start the service from first to last?
Ans: Zookeeper, HDFS, Mapredure2/Yarn, hive, spark...

Q 2. What are services you use for Authentication and Authorization
Ans: We use Kerberos for Authentication and ACL for Authorization.

Q 3. What is the size of your cluster and what are the services you use.
Ans: Cluster having 10 hosts
6 datanode, 2 Edge Node, 2 NameNode
6 hosts of 12 TB each
Blocksize=64 MB, Replication=3
12 TB * 6 Host = 72 TB
Cluster capacity in MB: 72 * 1000000 MB = 72,000,000 MB
Disk space needed per block: 64 MB per block * 3 = 192 MB storage per block
Total number of blocks: 72,000,000  / 192  = 375000 blocks
70% of total capacity i.e. 72 TB= 39 TB
Actual data = 13 TB
We can say 20-35 GB per day data. and keep last 12 months of data.

Note: Kindly correct it, if I am wrong. This is the only sketch

Q 4. What is the architecture of Hive?
Ans: https://selecthadoop.blogspot.in/search/label/Hive

Q 5. What are the Producer, Consumers, broker in Kafka?

Q 6. Execution of Hadoop Job.
Ans: 1. The client application submits a job to the resource manager.
2. The resource manager takes the job from the job queue and allocates it to an application master. It also manages and monitors resource allocations to each application master and container on the data nodes.
3. The application master divides the job into tasks and allocates it to each data node.
4. On each data node, a Node manager manages the containers in which the tasks run.
5. The application master will ask the resource manager to allocate more resource to particular containers, if necessary.
6. The application master will keep the resource manager informed as to the status of the jobs allocated to it, and the resource manager will keep the client application informed.

Q 7. What are the components of YARN.?

Q 8. What are your roles and responsibilities?
Ans: https://selecthadoop.blogspot.in/search/label/Daily%20Activities%20of%20Hadoop%20Admin

Q 9. What happens with the active namenode when any standby namenode become active?
Ans: Hi Hadoopers kindly help

Q 10.       What are the views in Amabari? Is the file views browse the local directory and Hadoop directory?
Ans: Ambari provides the UI for executing Hive query, pig scripts, transfer files from local to HDFS and vice versa. Yes, file views browse the local directory and Hadoop directory.

Q 11. How Hadoop save 100 MB of a file if your block size is 64 MB?
Ans: Hadoop save 100 MB of a file in 2 blocks, the first block size is of 64 MB and second block size is of 36 MB only.
Hadoop stores each file as a sequence of blocks; all blocks in a file except the last block are the same size.

Q 12. If we have 4 datanode and the replication factor is 3. So how we decommission the 2 datanodes from the cluster?
Ans:

Q 13. What are the steps of upgrade the Hadoop cluster? What are the changes to be made?
Ans:

Q 14: What happens when we do not give the snapshot name.
Ans: The snapshot name, which is an optional argument. When it is omitted, a default name is generated using a timestamp with the format "'s'yyyyMMdd-HHmmss.SSS", e.g. "s20130412-151029.033".

Q 15: In the kerberized Hadoop cluster, What are troubleshooting steps when any user is unable to login into the cluster.
Ans:

Q 16: Why do we use HDFS for applications having large data sets and not when there are lot of small files?
Ans: HDFS is more suitable for a large number of data sets in a single file as compared to small amount of data spread across multiple files. This is because Namenode is a very expensive high-performance system, so it is not prudent to occupy the space in the Namenode by an unnecessary amount of metadata that is generated for multiple small files. So, when there is a large amount of data in a single file, name node will occupy less space. Hence for getting optimized performance, HDFS supports large data sets instead of multiple small files.

Q 17: Web-UI shows that half of the datanodes are in decommissioning mode. What does that mean? Is it safe to remove those nodes from the network?
Ans: This means that namenode is trying to retrieve data from those datanodes by moving replicas to remain datanodes. There is a possibility that data can be lost if administrator removes those datanodes before decommissioning finished.Due to replication strategy, it is possible to lose some data due to datanodes removal en masse prior to completing the decommissioning process. Decommissioning refers to namenode trying to retrieve data from datanodes by moving replicas to remain datanodes.

Kafka Architecture

Apache Kafka is a distributed publish-subscribe messaging system and a robust queue that can handle a high volume of data and enables you t...