Use of Statistics in Data Science MCQs – Class 10 Data Science(419)
Practice this comprehensive Use of Statistics in Data Science MCQs collection covering all important topics and concepts from the chapter. It includes concept-based, application-based, and assertion-based questions so that no important question pattern is left uncovered.
Revise smarter, practice thoroughly, and build the confidence you need to score your best in the CBSE exam! 🚀
Q1. What is meant by a subset in Data Science?
a) The complete dataset
b) A smaller part of a larger dataset selected for analysis
c) A duplicate copy of the dataset
d) A mathematical formula
Q2. Why is subsetting useful in data analysis?
a) It increases unnecessary data
b) It removes the need for analysis
c) It helps focus on the required data
d) It changes all values in a dataset
Q3. A table contains 100 rows and 100 columns. A student selects the first 5 rows and 5 columns for analysis. What has the student created?
a) Frequency table
b) Two-way table
c) Subset
d) Relative frequency
Q4. A student selects the top 3 rows from a table containing 6 rows. Which type of subsetting is this?
a) Column-based
b) Data-based
c) Row-based
d) Relative
Q5. A dataset contains 50 columns, but only 4 are required for analysis. Which subsetting method is most appropriate?
a) Row-based
b) Column-based
c) Data-based
d) Median-based
Q6. Data-based subsetting selects data according to:
a) Specific data or conditions
b) The number of columns only
c) The average
d) The standard deviation
Q7. A two-way frequency table represents the observed frequency for:
a) One variable
b) Two variables
c) Three variables
d) No variables
Q8. In a two-way frequency table, rows generally indicate:
a) One category of a variable
b) The mean
c) The total only
d) Percentages only
Q9. In a two-way frequency table, the rows and columns generally represent:
a) Two different categories or variables
b) The same category repeated twice
c) Random unrelated numbers
d) Only column totals
Q10. What does an individual cell in a two-way frequency table represent?
a) Mean
b) Median
c) Frequency or number of people/data points
d) Standard deviation
Q11. A chocolate preference survey divides people into age groups and records whether they like chocolates. What type of table can represent this information?
a) One-way table
b) Two-way frequency table
c) Mean table
d) Median table
Q12. If a cell contains 6 in a frequency table, it means:
a) 6% of data
b) 6 data points/persons fall in that category
c) Mean is 6
d) Median is 6
Q13. A two-way relative frequency table differs from a two-way frequency table mainly because it uses:
a) Mean
b) Median
c) Percentages
d) Standard deviation
Q14. Two-way relative frequency tables make comparison easier because they use:
a) Raw values only
b) Percentages
c) Outliers
d) Variances
Q15. Mean is also commonly called:
a) Middle value
b) Simple average
c) Range
d) Deviation
Q16. Read the following statements carefully:
Statement I: Mean is a measure of central tendency.
Statement II: Mean is obtained by adding all values and dividing by the number of values.
a) Both Statement I and Statement II are true
b) Both Statement I and Statement II are false
c) Statement I is true, but Statement II is false
d) Statement I is false, but Statement II is true
Q17. Assertion: Median can be more effective than mean when a dataset contains outliers.
Reason: An outlier can cause the mean to deviate greatly from regular values.
Q18. The mean of a dataset is calculated by:
a) Multiplying all values
b) Adding all values and dividing by the number of values
c) Selecting the middle value
d) Subtracting the smallest value
Q19. In calculating mean, the values are:
a) Weighted equally
b) Always weighted differently
c) Ignored
d) Sorted first
Q20. The temperatures of five cities are 21°C, 13°C, 24°C, 15°C and 20°C. What is their mean?
a) 16.5°C
b) 17.5°C
c) 18.6°C
d) 20.5°C
Q21. Ravi, Juhi, Shweta and Kishan have heights 156 cm, 148 cm, 151 cm and 158 cm. What is their mean height?
a) 150.25 cm
b) 151.75 cm
c) 153.25 cm
d) 155 cm
Q22. If a dataset contains an odd number of observations, its median is:
a) Average of all values
b) Exact middle value after sorting
c) Largest value
d) Smallest value
Q23. Median is generally preferred over mean when:
a) There are outliers
b) There are no values
c) Dataset contains only percentages
d) Dataset is always small
Q24. Assertion: A low standard deviation can indicate that weather forecasts are more reliable.
Reason: Low standard deviation indicates less spread in the forecasted values.
a) Both Assertion and Reason are true, and Reason is the correct explanation of Assertion
b) Both Assertion and Reason are true, but Reason is not the correct explanation of Assertion
c) Assertion is true, but Reason is false
d) Assertion is false, but Reason is true
Q25. Assertion: Two datasets with the same mean must have the same variability.
Reason: Variability depends on how far the observations are from the mean.
a) Both Assertion and Reason are true, and Reason is the correct explanation of Assertion
b) Both Assertion and Reason are true, but Reason is not the correct explanation of Assertion
c) Assertion is true, but Reason is false
d) Assertion is false, but Reason is true
Q26. An unusually high blood-pressure reading caused by a device error is an example of:
a) Variable
b) Outlier
c) Subset
d) Frequency
Q27. Mean Absolute Deviation measures:
a) The middle value
b) The average distance of values from the mean
c) The largest value
d) The number of observations
Q28. The abbreviation MAD stands for:
a) Mean Average Data
b) Mean Absolute Deviation
c) Median Absolute Data
d) Mean Analysis Distribution
Q29. Why is the absolute value used in MAD?
a) To ignore the negative sign
b) To increase the mean
c) To calculate the median
d) To sort the dataset
Q30. Standard deviation measures:
a) How spread out the numbers are
b) Only the mean
c) Only the median
d) Number of categories
Q31. Standard deviation represents how much data is spread around:
a) Median
b) Mean
c) Maximum
d) Minimum
Q32. The average of squared differences is called:
a) Median
b) Variance
c) MAD
d) Frequency
Q33. Standard deviation is obtained by:
a) Taking square root of variance
b) Squaring variance
c) Adding variance to mean
d) Dividing variance by median
Q34. For the dataset 1, 2, 3, 5, 8, the mean is:
a) 2.8
b) 3.8
c) 4.8
d) 5.8
Q35. For the values 1, 2, 3, 5 and 8 (mean = 3.8), what is the variance?
a) 4.16
b) 5.16
c) 6.04
d) 7.16
Q36. A teacher wants to determine whether students are performing at a similar level. Which measure can help examine the spread of scores?
a) Standard deviation
b) Subsetting
c) Median only
d) Relative frequency only
Q37. According to the chapter, a low standard deviation in weather forecasting indicates:
a) Less reliable forecast
b) More reliable forecast
c) No forecast
d) Higher median
Q38. If five families have an equal share value of 6, the total number of snap cubes is:
a) 11
b) 20
c) 30
d) 36
Q39. Two groups can have the same mean but:
a) Never have different distributions
b) Have different distributions and variability
c) Never contain five observations
d) Always have the same median
Q40. Read the following statements carefully:
Statement I: Standard deviation measures how spread out data is.
Statement II: Standard deviation is obtained by taking the square root of variance.
a) Both Statement I and Statement II are true
b) Both Statement I and Statement II are false
c) Statement I is true, but Statement II is false
d) Statement I is false, but Statement II is true