Understanding data analysis

 Understanding data analysis

The 21st century is the century of information. We are living in the age of information,

which means that almost every aspect of our daily life is generating data. Not only this, but

business operations, government operations, and social posts are also generating huge data.

This data is accumulating day by day due to data being continually generated from

business, government, scientific, engineering, health, social, climate, and environmental

activities. In all these domains of decision-making, we need a systematic, generalized,

effective, and flexible system for the analytical and scientific process so that we can gain

insights into the data that is being generated.

In today's smart world, data analysis offers an effective decision-making process for

business and government operations. Data analysis is the activity of inspecting, preprocessing, exploring, describing, and visualizing the given dataset. The main objective of

the data analysis process is to discover the required information for decision-making. Data

analysis offers multiple approaches, tools, and techniques, all of which can be applied to

diverse domains such as business, social science, and fundamental science.

Let's look at some of the core fundamental data analysis libraries of the Python ecosystem:

NumPy: This is a short form of numerical Python. It is the most powerful

scientific library available in Python for handling multidimensional arrays,

matrices, and methods in order to compute mathematics efficiently.

SciPy: This is also a powerful scientific computing library for performing

scientific, mathematical, and engineering operations.

Pandas: This is a data exploration and manipulation library that offers tabular

data structures such as DataFrames and various methods for data analysis and

manipulation.

Scikit-learn: This stands for "Scientific Toolkit for Machine learning". It is a

machine learning library that offers a variety of supervised and unsupervised

algorithms, such as regression, classification, dimensionality reduction, cluster

analysis, and anomaly detection.

Matplotlib: This is a core data visualization library and is the base library for all

other visualization libraries in Python. It offers 2D and 3D plots, graphs, charts,

and figures for data exploration. It runs on top of NumPy and SciPy.

Seaborn: This is based on Matplotlib and offers easy to draw, high-level,

interactive, and more organized plots.

Plotly: Plotly is a data visualization library. It offers high quality and interactive

graphs, such as scatter charts, line charts, bar charts, histograms, boxplots,

heatmaps, and subplots.





The standard process of data analysis

Data analysis refers to investigating the data, finding meaningful insights from it, and

drawing conclusions. The main goal of this process is to collect, filter, clean, transform,

explore, describe, visualize, and communicate the insights from this data to discover

decision-making information. Generally, the data analysis process is comprised of the

following phases:

1. Collecting Data: Collect and gather data from several sources.

2. Preprocessing Data: Filter, clean, and transform the data into the required

format.

3. Analyzing and Finding Insights: Explore, describe, and visualize the data and

find insights and conclusions.

4. Insights Interpretations: Understand the insights and find the impact each

variable has on the system.

5. Storytelling: Communicate your results in the form of a story so that a layman

can understand them.






The KDD process

The KDD acronym stands for knowledge discovery from data or Knowledge Discovery in Databases. Many people treat KDD as one synonym for data mining. Data mining is referred to as the knowledge discovery process of interesting patterns. The main objective of KDD is to extract or discover hidden interesting patterns from large databases, data warehouses, and other web and information repositories. The KDD process has seven major phases:

1. Data Cleaning: In this first phase, data is preprocessed. Here, noise is removed, missing values are handled, and outliers are detected.

2. Data Integration: In this phase, data from different sources is combined and integrated together using data migration and ETL tools.

3. Data Selection: In this phase, relevant data for the analysis task is recollected.
Data Transformation: In this phase, data is engineered in the required appropriate form for analysis.

5. Data Mining: In this phase, data mining techniques are used to discover useful and unknown patterns.

6. Pattern Evaluation: In this phase, the extracted patterns are evaluated.

7. Knowledge Presentation: After pattern evaluation, the extracted knowledge needs to be visualized and presented to business people for decision-making purposes.




SEMMA

The SEMMA acronym's full form is Sample, Explore, Modify, Model, and Assess. This

sequential data mining process is developed by SAS. The SEMMA process has five major

phases:

1. Sample: In this phase, we identify different databases and merge them. After this, we select the data sample that's sufficient for the modeling process.

2. Explore: In this phase, we understand the data, discover the relationships among variables, visualize the data, and get initial interpretations.

3. Modify: In this phase, data is prepared for modeling. This phase involves dealing with missing values, detecting outliers, transforming features, and creating new additional features.

4. Model: In this phase, the main concern is selecting and applying different modeling techniques, such as linear and logistic regression, backpropagation networks, KNN, support vector machines, decision trees, and Random Forest.

5. Assess: In this last phase, the predictive models that have been developed are evaluated using performance evaluation measures

The preceding diagram shows the steps involved in the SEMMA process. SEMMA emphasizes model building and assessment. Now, let's discuss the CRISP-DM process.



CRISP-DM

CRISP-DM's full form is CRoss-InduStry Process for Data Mining. CRISP-DM is a welldefined, well-structured, and well-proven process for machine learning, data mining, and

business intelligence projects. It is a robust, flexible, cyclic, useful, and practical approach to

solving business problems. The process discovers hidden valuable information or patterns

from several databases. The CRISP-DM process has six major phases:

1. Business Understanding: In this first phase, the main objective is to understand

the business scenario and requirements for designing an analytical goal and

initial action plan.

2. Data Understanding: In this phase, the main objective is to understand the data

and its collection process, perform data quality checks, and gain initial insights.

3. Data Preparation: In this phase, the main objective is to prepare analytics-ready

data. This involves handling missing values, outlier detection and handling,

normalizing data, and feature engineering. This phase is the most timeconsuming for data scientists/analysts.

4. Modeling: This is the most exciting phase of the whole process since this is

where you design the model for prediction purposes. First, the analyst needs to

decide on the modeling technique and develop models based on data.

5. Evaluation: Once the model has been developed, it's time to assess and test the

model's performance on validation and test data using model evaluation

measures such as MSE, RMSE, R-Square for regression and accuracy, precision,

recall, and the F1-measure.

6. Deployment: In this final phase, the model that was chosen in the previous step

will be deployed to the production environment. This requires a team effort from

data scientists, software developers, DevOps experts, and business professionals.


The following diagram shows the full cycle of the CRISP-DM process:






The standard process focuses on discovering insights and making interpretations in the
form of a story, while KDD focuses on data-driven pattern discovery and visualizing this.
SEMMA majorly focuses on model building tasks, while CRISP-DM focuses on business
understanding and deployment. Now that we know about some of the processes
surrounding data analysis, let's compare data analysis and data science to find out how
they are related, as well as what makes them different from one other.


Comparing data analysis and data science

Data analysis is the process in which data is explored in order to discover patterns that help
us make business decisions. It is one of the subdomains of data science. Data analysis
methods and tools are widely utilized in several business domains by business analysts,
data scientists, and researchers. Its main objective is to improve productivity and
profits. Data analysis extracts and queries data from different sources, performs exploratory
data analysis, visualizes data, prepares reports, and presents it to the business decisionmaking authorities.
On the other hand, data science is an interdisciplinary area that uses a scientific approach to
extract insights from structured and unstructured data. Data science is a union of all terms,
including data analytics, data mining, machine learning, and other related domains. Data
science is not only limited to exploratory data analysis and is used for developing models
and prediction algorithms such as stock price, weather, disease, fraud forecasts, and
recommendations such as movie, book, and music recommendations.



The roles of data analysts and data scientists

A data analyst collects, filters, processes, and applies the required statistical concepts to capture patterns, trends, and insights from data and prepare reports for making decisions.

The main objective of the data analyst is to help companies solve business problems using discovered patterns and trends. The data analyst also assesses the quality of the data and handles the issues concerning data acquisition. A data analyst should be proficient in writing SQL queries, finding patterns, using visualization tools, and using reporting tools Microsoft Power BI, IBM Cognos, Tableau, QlikView, Oracle BI, and more. Data scientists are more technical and mathematical than data analysts. Data scientists are research- and academic-oriented, whereas data analysts are more application-oriented. Data scientists are expected to predict a future event, whereas data analysts extract significant insights out of data. Data scientists develop their own questions, while data analysts find answers to given questions. Finally, data scientists focus on what is going to happen, whereas data analysts focus on what has happened so far. We can summarize these two roles using the following.


Features Data Scientist Data Analyst
Background Predict future events and scenarios based on data Discover meaningful insights from the data.
Role Formulate questions that can profit the businessSolve the business questions to make decisions.

Type of data Work on both structured and unstructured data Only work on structured data
Programming Advanced programming Basic programming
SkillsetKnowledge of statistics, machine learning algorithms, NLP, and deep learningKnowledge of statistics, SQL, and data visualization
Tools R, Python, SAS, Hadoop, Spark, TensorFlow, and KerasExcel, SQL, R, Tableau, and QlikView


Now that we know what defines a data analyst and data scientist, as well as how they are different from each other, let's have a look at the various skills that you would need to become one of them.




MPLS TROUBLESHOOTING TIPS FOR CISCO AND JUNIPER

MPLS TROUBLESHOOTING TIPS FOR CISCO AND JUNIPER


#mpls #cisco #juniper #troubleshooting #huawei #copy #tutorial

Basic MPLS Troubleshooting Tips

1. Verifying MPLS Configuration:

  • Cisco:
    • Use show mpls interfaces to verify that MPLS is enabled on the correct interfaces.
    • Check show mpls ldp neighbor to ensure that Label Distribution Protocol (LDP) neighbors are discovered, and that the session is up.
  • Juniper:
    • Use show mpls interface to check MPLS status on interfaces.
    • Utilize show mpls ldp session to confirm LDP neighbor sessions.

2. Checking Label Switch Paths (LSP):

  • Cisco:
    • Use show mpls ldp bindings to display local and remote label bindings.
    • show mpls forwarding-table helps to inspect the labels being forwarded and their corresponding next-hops.
  • Juniper:
    • Use show mpls lsp extensive to get detailed information about the LSPs.
    • show route table mpls.0 to view the label-switched routes.

3. Ensuring Proper Route Distribution:

  • Cisco:
    • Verify routing protocols are correctly redistributing routes with show ip route and show ip protocols.
    • Ensure that MPLS labels are being properly assigned by checking show mpls forwarding-table.
  • Juniper:
    • Check routing information with show route forwarding-table family mpls.
    • Ensure correct route redistribution settings with show route protocol.

4. Troubleshooting MPLS VPNs:

  • Cisco:
    • For issues with VRF (Virtual Routing and Forwarding), use show ip vrf and show ip route vrf [vrf-name].
    • Verify MPLS VPN label distribution and path information using show mpls forwarding-table vrf [vrf-name].
  • Juniper:
    • Check VRFs using show route table [vrf-name].inet.0.
    • Look at the VPN labels with show route table [vrf-name].inet.0 detail.

5. Utilizing Cisco debug and Juniper traceoptions:

  • Cisco:
    • In-depth troubleshooting can be performed by enabling debugging: debug mpls ldp for LDP-related issues or debug mpls traffic-eng for traffic engineering problems.
  • Juniper:
    • Use traceoptions under the MPLS or routing protocol configuration to capture more detailed logs for troubleshooting.

6. Common Pitfalls and Checks:

  • Both Cisco and Juniper:
    • Ensure there are no MTU mismatches across MPLS-enabled interfaces, as this can disrupt proper LSP formation.
    • Regularly check for software or firmware updates that address known bugs or add enhancements to MPLS features.

Prepare Data for Exploration : Weekly challenge 2

 Prepare Data for Exploration : Weekly challenge 2


cybersecurity, cisco, coursera, quiz, solution, availability, security,cyber, ops, dev, correct, answer, privacy

#cybersecurity #coursera #quiz #solution #network


Coursera google professional data analytical course 3 weekly challange 2 answers



1.

Question 1

Fill in the blank: A preference in favor of or against a person, group of people, or thing is called _____. It is an error in data analytics that can systematically skew results in a certain direction.

1 / 1 point

data interoperability

data collection

data bias

data anonymization

Correct

2.

Question 2

A data analyst studies the sales data obtained after each marketing campaign to determine the effectiveness of the campaign. When the findings are ambiguous, the analyst chooses to interpret the results positively. What type of bias does this represent?

0 / 1 point

Sampling

Confirmation

Interpretation

Observer

Incorrect

3.

Question 3

A data analyst reviews a dataset. They conclude that the data is inaccurate and incomplete in some places. They also confirm that the data is biased. What type of data does this describe?

1 / 1 point

Informed data

Good data

Open data

Unreliable data

Correct

4.

Question 4

Fill in the blank: Data _____ refers to well-founded standards of right and wrong that dictate how data is collected, shared, and used.

1 / 1 point

ethics

anonymization

credibility

privacy

Correct

5.

Question 5

Ownership is a key issue in data ethics. Who owns data?

0 / 1 point

The organization that invests time and money collecting, processing, and analyzing the data

The law enforcement agencies that enforce data protection laws

The government that passes data-protection legislation

The individual who originally generates the data

Incorrect

6.

Question 6

In data ethics, which of the following rights is included in data privacy? Select all that apply.

0 / 1 point

The right to inspect the data.

The right to meet the people who will be working on the data.

This should not be selected

The right to update the data.

The right to correct the data.

7.

Question 7

Which of the following are commonly used methods for anonymizing data? Select all that apply.

0.5 / 1 point

Deleting

This should not be selected

Blanking

Masking

Correct

Hashing

Correct

8.

Question 8

The government of a large city collects data on the quality of the city’s infrastructure. Any business, nonprofit organization, or citizen can access the government’s databases and re-use or re-distribute the data. Is this an example of open data?

1 / 1 point

Yes

No

Correct

 

 

1.

Question 1

Fill in the blank: Data _____ is a preference in favor of or against a person, group of people, or thing. In data analytics, it can systematically skew results in a certain direction.

0 / 1 point

collection

anonymization

interoperability

bias

Incorrect

Please review the video on errors that can skew results.

2.

Question 2

A university surveys its student-athletes about their experience in college sports. The survey only includes student-athletes with scholarships. What type of bias is this an example of?

1 / 1 point

Confirmation bias

Sampling bias

Observer bias

Interpretation bias

Correct

3.

Question 3

Which of the following “C’s” describe qualities of good data? Select all that apply.

1 / 1 point

Cited

Correct

Current

Correct

Comprehensive

Correct

Consequential

4.

Question 4

In data ethics, what gives an individual the right to know why their data is collected and how it will be used?

0 / 1 point

Privacy

Consent

Anonymization

Credibility

Incorrect

Please review the video on well-founded standards of right and wrong.

5.

Question 5

An individual who provides their data has the right to know and understand all of the data-processing activities and algorithms used on that data. This is called transaction transparency.

1 / 1 point

True

False

Correct

6.

Question 6

A data analyst is working with sensitive data from their client. Because the data is sensitive, the analyst reminds the client of their data privacy rights. What rights do these include? Select all that apply.

1 / 1 point

The right to update the data.

Correct

The right to inspect the data.

Correct

The right to meet the people who will work on the data.

The right to correct the data.

Correct

7.

Question 7

Data anonymization applies to both text and images.

1 / 1 point

True

False

Correct

8.

Question 8

A company decides to allow people access to its databases for the purposes of re-using and re-distributing data. In order to facilitate this, they use a common terminology and format. This is an example of what?

1 / 1 point

Courtesy

Interoperability

Convenience

Standardization

Correct

 

 

1.

Question 1

A clinic surveys a group of male and female patients about their experience with physical therapy. The survey does not include people with disabilities. Is the survey data biased?

1 / 1 point

Yes

No

Correct

2.

Question 2

Which of the following are types of data bias often encountered in data analytics? Select all that apply.

1 / 1 point

Educational bias

Confirmation bias

Correct

Interpretation bias

Correct

Observer bias

Correct

3.

Question 3

In general, the usefulness of data decreases as time passes.

1 / 1 point

True

False

Correct

4.

Question 4

In data ethics, what gives an individual the right to know why their data is collected and how it will be used?

0 / 1 point

Consent

Anonymization

Credibility

Privacy

5.

Question 5

Transactional transparency is a fundamental right for an individual who provides their data. Which of the following is included in this right? Select all that apply.

0.75 / 1 point

Knowing the data-processing activities to be used on that data.

Correct

Meeting the people who will be working on the data.

Knowing for how long the data will be used.

Understanding the algorithms to be used on the data.

Correct

6.

Question 6

What is data privacy?

0 / 1 point

Providing free access, usage, and sharing of data

Preserving a data subject’s information and activity for all data transactions

Searching for or interpreting supporting information

Applying standards that preserve the consistency in how data is collected, shared, and used

Incorrect

7.

Question 7

Why would a company routinely use a data anonymizer when working with their users’ data?

1 / 1 point

To make it easier to recognize what data corresponds to which individual

To keep its users’ data consistent by removing any identifying information

To eliminate data bias caused by sensitive information

To protect its users’ private and sensitive data by removing any identifying information

Correct

8.

Question 8

Fill in the blank: When a company or organization allows their data to be re-used and re-distributed that data is considered to be _____ .

1 / 1 point

free

allowable

closed

open

Correct

 





Question 1

Fill in the blank: A preference in favor of or against a person, group of people, or thing is called _____. It is an error in data analytics that can systematically skew results in a certain direction.

0 / 1 point
Incorrect
Question 2

A researcher believes that playing music to plants will increase flower production. They set up an experiment and collect the data. Although the findings are inconclusive, they choose to interpret the data in a way that supports their desired result. What type of bias does this represent?

0 / 1 point
Incorrect
Question 3

A data analyst reviews a dataset. They confirm that the data is comprehensive, cited, and current. What type of data does this describe?

1 / 1 point
Correct
Question 4

If a company uses your personal data as part of a financial transaction, you should be made aware of the nature and scale of the transaction. What concept of data ethics does this refer to?

1 / 1 point
Correct
Question 5

An individual who provides their data has the right to know and understand all of the data-processing activities and algorithms used on that data. This concept refers to which aspect of data ethics?

1 / 1 point
Correct
Question 6

An employer accesses an employee’s credit report without their consent. This is not a violation of the employee’s privacy because they work at the company.

1 / 1 point
Correct
Question 7

Which of the following are commonly used methods for anonymizing data? Select all that apply.

0.75 / 1 point
Correct
Correct
You didn’t select all the correct answers
Question 8

A government agency allows any business, nonprofit organization, or citizen to access the government’s databases and re-use or re-distribute the data. What type of data is this an example of?

1 / 1 point

Correct



Question 1

Which of the following best describes data bias?

1 / 1 point
Correct
Question 2

Which of the following are types of data bias often encountered in data analytics? Select all that apply.

1 / 1 point
Correct
Correct
Correct
Question 3

Which of the following “C’s” describe qualities of good data? Select all that apply.

1 / 1 point
Correct
Correct
Correct
Question 4

Before completing a survey, a customer wants to learn more about how a company will use their data. They want to know why their data is being collected, how it will be used, and how long it will be stored. What data ethics concept does this describe?

0 / 1 point
Incorrect


Question 5

Transactional transparency is a fundamental right for an individual who provides their data. Which of the following is included in this right? Select all that apply.

1 / 1 point
Correct
Correct
Question 6

The right to inspect, update, or correct your own data is part of which aspect of data ethics?

1 / 1 point
Correct
Question 7

Data anonymization applies to both text and images.

1 / 1 point
Correct
Question 8

A key aspect of open data is free access to people’s personal information.

1 / 1 point
Correct

Featured Post

Day 41 — BGP Confederations: Sub-AS Design, External View and Migration

1. Opening Confederations are another way to scale BGP inside a large administrative domain. They divide the domain into member autonomous systems while presenting a single confederation identifier to external peers. They are powerful, but their operational model is more complex than simply 'using private ASNs inside.' The engineering goal is not to memorize another BGP command. It is to understand what information each speaker is allowed to propagate, what path information can be hidden, and what failure domain is created by the chosen control-plane architecture . 2. Concept and standards behavior RFC 5065 defines AS_CONFED_SEQUENCE and AS_CONFED_SET and how member-AS relationships are represented. Confederation external sessions have eBGP-like properties inside the confederation, while the confederation is presented externally as one AS. Modern guidance must also account for the fact that RFC 9774 prohibits new origination of AS_SET/AS_CONFED_SET in ordinary aggregation c...