> For the complete documentation index, see [llms.txt](https://tony-ng-1.gitbook.io/what-is-ml/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://tony-ng-1.gitbook.io/what-is-ml/ml-scenario/terminologies.md).

# Terminologies

## Terminologies

Before we fly, let's crawl through some commonly used terminologies.

![you can do it](https://media3.giphy.com/media/l4Epk31YxoTq7KYr6/giphy.gif)

### Input variable (x)

* The information used to **learn**
* There can be **more than 1** variable
* For example, **"Gender", "Age"**, and **"Salary"** in the credit-card example.

### Output variable (y)

* The desired **output** after learning
* There can **only be** **1** output variable
* For discrete, this can either be **defaulting or not defaulting** in credit-card&#x20;
* For continuous, this can be predicting the **amount of credit** to give

### Target function (f)&#x20;

* 'f' here refers to the **ideal** hypothesis&#x20;
* A hypothesis **maps** a sample to its various y values. With this, the actual training examples are **'generated'**&#x20;
* **f : x → y**, means eating some x variables, and spit out y variables
* f is **unknown**

![A function](https://media1.giphy.com/media/xTiQyrlPe79FJF60iA/giphy.gif)

I like analogies, especially this one from 3Blue1Brown (I believe). A function merely **eats up data**, and **spits out data**. Different function transforms data differently, in which can be interchangeably referred to as *hypothesis*.&#x20;

Also, I understand that point 2 can be extremely confusing (at least for me, but will further elaborate it to my best in the next topic), but in short it is that the **'real**' set of output variable is generated with the target function f.&#x20;

### Data (x, y)

* Generated from the **target function f(x)**
* A combination of both **input (x)** and **output (y)** variables

$$
(x\_1,y\_1),(x\_2,y\_2),...,(x\_N,y\_N)
$$

* Each instance is represented by the **row number N**

### Hypothesis (g)

* We can have **multiple** hypotheses in our **hypothesis set (H)**.
* However, there can only be **one** (and only one) we select to be called g
* Hence, $$g \in H$$, or **g is in H**
* g is the **best** hypothesis approximating the f, $$g \approx f$$&#x20;

To put it all together,

![Model of ML](https://1758022264-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-Lxakp1GCvqByN-Lj-sA%2F-Ly8_sZlJxO92hFE8ASp%2F-Ly8a7HUGJG14FtTjhSk%2Fimage.png?alt=media\&token=3d3880c8-ebe5-40dc-bf2b-d544fd70f0b9)
