<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Meta Learning on eree's blog</title><link>https://blog.ereebay.me/en/tags/meta-learning/</link><description>Recent content in Meta Learning on eree's blog</description><generator>Hugo</generator><language>en</language><copyright>2020-2026 eree&amp;rsquo;s blog</copyright><lastBuildDate>Fri, 10 Jan 2020 11:31:05 +0800</lastBuildDate><atom:link href="https://blog.ereebay.me/en/tags/meta-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>CS330 Lecture 1&amp;2 Study Notes (Incomplete)</title><link>https://blog.ereebay.me/en/posts/cs330-1/</link><pubDate>Fri, 10 Jan 2020 11:31:05 +0800</pubDate><guid>https://blog.ereebay.me/en/posts/cs330-1/</guid><description>&lt;h1 id="cs330-lecture-12-notes"&gt;CS330 lecture 1&amp;amp;2 notes&lt;/h1&gt;
&lt;h2 id="informal-problem-definitions"&gt;Informal Problem Definitions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;The multi-task learning problem: Learn all of the tasks more quickly or more proficiently than learning them independently.&lt;/li&gt;
&lt;li&gt;The meta-learning problem: Given data/experience on previous tasks, learn a new task more quickly and/or more proficiently.&lt;/li&gt;
&lt;/ul&gt;
&lt;!-- more --&gt;
&lt;h2 id="multi-task-learning-basics"&gt;Multi-Task Learning Basics&lt;/h2&gt;
&lt;p&gt;Traditional single-task learning:&lt;/p&gt;
&lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mtable rowspacing="0.16em" columnalign="left" columnspacing="1em"&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="false"&gt;&lt;mrow&gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mo fence="true"&gt;{&lt;/mo&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;y&lt;/mi&gt;&lt;msub&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence="true"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;mtr&gt;&lt;mtd&gt;&lt;mstyle scriptlevel="0" displaystyle="false"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mi&gt;min&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mstyle&gt;&lt;/mtd&gt;&lt;/mtr&gt;&lt;/mtable&gt;&lt;annotation encoding="application/x-tex"&gt;
\begin{array}{l}{\mathscr{D}=\left\{(\mathbf{x}, \mathbf{y})_{k}\right\}} \\ {\min _{\theta} \mathscr{L}(\theta, \mathscr{D})}\end{array}
&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;p&gt;Typical loss: negative log likelihood&lt;/p&gt;
&lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="double-struck"&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;mi mathvariant="script"&gt;D&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence="true"&gt;[&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;y&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo fence="true"&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;
\mathscr{L}(\theta, \mathscr{D})=-\mathbb{E}_{(x, y) \sim \mathscr{D}}\left[\log f_{\theta}(\mathbf{y} | \mathbf{x})\right]
&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;h3 id="whats-a-task"&gt;What&amp;rsquo;s a task?&lt;/h3&gt;
&lt;p&gt;A task: &lt;span class="katex"&gt;&lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;≜&lt;/mo&gt;&lt;mrow&gt;&lt;mo fence="true"&gt;{&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo stretchy="false"&gt;(&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;y&lt;/mi&gt;&lt;mi mathvariant="normal"&gt;∣&lt;/mi&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mo stretchy="false"&gt;)&lt;/mo&gt;&lt;mo separator="true"&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="script"&gt;L&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence="true"&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;\mathscr{T}_{i} \triangleq\left\{p_{i}(\mathbf{x}), p_{i}(\mathbf{y} | \mathbf{x}), \mathscr{L}_{i}\right\}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/p&gt;</description><content:encoded><![CDATA[<h1 id="cs330-lecture-12-notes">CS330 lecture 1&amp;2 notes</h1>
<h2 id="informal-problem-definitions">Informal Problem Definitions</h2>
<ul>
<li>The multi-task learning problem: Learn all of the tasks more quickly or more proficiently than learning them independently.</li>
<li>The meta-learning problem: Given data/experience on previous tasks, learn a new task more quickly and/or more proficiently.</li>
</ul>
<!-- more -->
<h2 id="multi-task-learning-basics">Multi-Task Learning Basics</h2>
<p>Traditional single-task learning:</p>
<span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mtable rowspacing="0.16em" columnalign="left" columnspacing="1em"><mtr><mtd><mstyle scriptlevel="0" displaystyle="false"><mrow><mi mathvariant="script">D</mi><mo>=</mo><mrow><mo fence="true">{</mo><mo stretchy="false">(</mo><mi mathvariant="bold">x</mi><mo separator="true">,</mo><mi mathvariant="bold">y</mi><msub><mo stretchy="false">)</mo><mi>k</mi></msub><mo fence="true">}</mo></mrow></mrow></mstyle></mtd></mtr><mtr><mtd><mstyle scriptlevel="0" displaystyle="false"><mrow><msub><mrow><mi>min</mi><mo>⁡</mo></mrow><mi>θ</mi></msub><mi mathvariant="script">L</mi><mo stretchy="false">(</mo><mi>θ</mi><mo separator="true">,</mo><mi mathvariant="script">D</mi><mo stretchy="false">)</mo></mrow></mstyle></mtd></mtr></mtable><annotation encoding="application/x-tex">
\begin{array}{l}{\mathscr{D}=\left\{(\mathbf{x}, \mathbf{y})_{k}\right\}} \\ {\min _{\theta} \mathscr{L}(\theta, \mathscr{D})}\end{array}
</annotation></semantics></math></span><p>Typical loss: negative log likelihood</p>
<span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">L</mi><mo stretchy="false">(</mo><mi>θ</mi><mo separator="true">,</mo><mi mathvariant="script">D</mi><mo stretchy="false">)</mo><mo>=</mo><mo>−</mo><msub><mi mathvariant="double-struck">E</mi><mrow><mo stretchy="false">(</mo><mi>x</mi><mo separator="true">,</mo><mi>y</mi><mo stretchy="false">)</mo><mo>∼</mo><mi mathvariant="script">D</mi></mrow></msub><mrow><mo fence="true">[</mo><mi>log</mi><mo>⁡</mo><msub><mi>f</mi><mi>θ</mi></msub><mo stretchy="false">(</mo><mi mathvariant="bold">y</mi><mi mathvariant="normal">∣</mi><mi mathvariant="bold">x</mi><mo stretchy="false">)</mo><mo fence="true">]</mo></mrow></mrow><annotation encoding="application/x-tex">
\mathscr{L}(\theta, \mathscr{D})=-\mathbb{E}_{(x, y) \sim \mathscr{D}}\left[\log f_{\theta}(\mathbf{y} | \mathbf{x})\right]
</annotation></semantics></math></span><h3 id="whats-a-task">What&rsquo;s a task?</h3>
<p>A task: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="script">T</mi><mi>i</mi></msub><mo>≜</mo><mrow><mo fence="true">{</mo><msub><mi>p</mi><mi>i</mi></msub><mo stretchy="false">(</mo><mi mathvariant="bold">x</mi><mo stretchy="false">)</mo><mo separator="true">,</mo><msub><mi>p</mi><mi>i</mi></msub><mo stretchy="false">(</mo><mi mathvariant="bold">y</mi><mi mathvariant="normal">∣</mi><mi mathvariant="bold">x</mi><mo stretchy="false">)</mo><mo separator="true">,</mo><msub><mi mathvariant="script">L</mi><mi>i</mi></msub><mo fence="true">}</mo></mrow></mrow><annotation encoding="application/x-tex">\mathscr{T}_{i} \triangleq\left\{p_{i}(\mathbf{x}), p_{i}(\mathbf{y} | \mathbf{x}), \mathscr{L}_{i}\right\}</annotation></semantics></math></span></p>
<p>data generating distributions</p>
<p>Here a task is defined as the distribution over data samples, the distribution over data labels, and a loss function.</p>
<p>Corresponding datasets: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msubsup><mi mathvariant="script">D</mi><mi>i</mi><mrow><mi>t</mi><mi>r</mi></mrow></msubsup></mrow><annotation encoding="application/x-tex">\mathscr{D}_{i}^{tr}</annotation></semantics></math></span> training set, <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msubsup><mi mathvariant="script">D</mi><mi>i</mi><mrow><mi>t</mi><mi>s</mi><mi>t</mi></mrow></msubsup></mrow><annotation encoding="application/x-tex">\mathscr{D}_{i}^{t s t}</annotation></semantics></math></span> test set.</p>
<p>Usually <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="script">D</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">\mathscr{D}_{i}</annotation></semantics></math></span> denotes the training set.</p>
<p>Multi-task classification: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="script">L</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">\mathscr{L}_{i}</annotation></semantics></math></span> same across all tasks. E.g., in handwritten character recognition across different languages, the form of the loss function may be the same.</p>
<p>Multi-label learning: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="script">L</mi><mi>i</mi></msub><mo separator="true">,</mo><msub><mi>p</mi><mi>i</mi></msub><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\mathscr{L}_{i}, {p}_{i}(x)</annotation></semantics></math></span> same across all tasks. E.g., in the CelebA multi-label recognition task, the samples and the loss function are identical.</p>
<p>The loss function may vary across tasks in the following cases:</p>
<ul>
<li>mixed discrete, continuous labels across tasks</li>
<li>caring more about one task than another (i.e., different weights for different tasks?)</li>
</ul>
<h3 id="conditioning-on-the-task">Conditioning on the task</h3>
<p>The multi-task learning problem requires introducing a task descriptor as a variable that describes the task; the question is how to design this variable.</p>
<p>Assume <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>z</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">{z}_{i}</annotation></semantics></math></span> is the task index. The most straightforward approach is multiplicative gating, which effectively trains each task in the multi-task setting with its own separate network, without sharing parameters.</p>
<p>The other extreme is to directly concat <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>z</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">z_i</annotation></semantics></math></span>, in which case all parameters are shared except those that come after the input <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>z</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">z_i</annotation></semantics></math></span>.</p>
<p>Yet another idea is to split <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi></mrow><annotation encoding="application/x-tex">\theta</annotation></semantics></math></span> into shared parameters <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi>θ</mi><mrow><mi>s</mi><mi>h</mi></mrow></msup></mrow><annotation encoding="application/x-tex">\theta^{sh}</annotation></semantics></math></span> and task-specific parameters <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi>θ</mi><mi>i</mi></msup></mrow><annotation encoding="application/x-tex">\theta^i</annotation></semantics></math></span> — i.e., shared and non-shared parameters.</p>
<p>The optimization objective then becomes</p>
<span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mrow><mi>min</mi><mo>⁡</mo></mrow><mrow><msup><mi>θ</mi><mrow><mi>s</mi><mi>h</mi></mrow></msup><mo separator="true">,</mo><msup><mi>θ</mi><mn>1</mn></msup><mo separator="true">,</mo><mo>…</mo><mo separator="true">,</mo><msup><mi>θ</mi><mi>T</mi></msup></mrow></msub><msubsup><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></msubsup><msub><mi mathvariant="script">L</mi><mi>i</mi></msub><mrow><mo fence="true">(</mo><mrow><mo fence="true">{</mo><msup><mi>θ</mi><mrow><mi>s</mi><mi>h</mi></mrow></msup><mo separator="true">,</mo><msup><mi>θ</mi><mi>i</mi></msup><mo fence="true">}</mo></mrow><mo separator="true">,</mo><msub><mi mathvariant="script">D</mi><mi>i</mi></msub><mo fence="true">)</mo></mrow></mrow><annotation encoding="application/x-tex">
\min _{\theta^{s h}, \theta^{1}, \ldots, \theta^{T}} \sum_{i=1}^{T} \mathscr{L}_{i}\left(\left\{\theta^{s h}, \theta^{i}\right\}, \mathscr{D}_{i}\right)
</annotation></semantics></math></span><p>The problem then becomes which parameters to share and when.</p>
<h4 id="common-choices">Common Choices</h4>
<p>The common choices are mainly concatenation and addition — the figures make them clear at a glance.</p>
<ol>
<li>Concatenation-based conditioning</li>
</ol>
<p><img alt="cs330-1-1.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-1.png"></p>
<ol start="2">
<li>Additive conditioning</li>
</ol>
<p><img alt="cs330-1-2.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-2.png"></p>
<p>In fact, the two are equivalent.</p>
<p><img alt="cs330-1-3.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-3.png"></p>
<ol start="3">
<li>Multi-head architecture</li>
</ol>
<p><img alt="cs330-1-4.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-4.png"></p>
<ol start="4">
<li>Multiplicative conditioning</li>
</ol>
<p><img alt="cs330-1-5.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-5.png"></p>
<p>The multiplicative approach offers:</p>
<ul>
<li>stronger expressive power</li>
<li>multiplication gating for regression tasks</li>
<li>better generalization across independent networks and heads</li>
</ul>
<h4 id="complex-choices">Complex Choices</h4>
<p>There are also many other more complex choices.</p>
<p><img alt="cs330-1-6.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-6.png"></p>
<p>But where the design inspiration comes from is just like choosing the hyperparameters of a neural network:</p>
<ul>
<li>different problems are independent of one another</li>
<li>for any specific problem, it mostly relies on the designer&rsquo;s intuition and background knowledge</li>
<li>current approaches are more art than science</li>
</ul>
<h3 id="optimizing-the-objective">Optimizing the objective</h3>
<p>Objective: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mrow><mi>min</mi><mo>⁡</mo></mrow><mi>θ</mi></msub><msubsup><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></msubsup><msub><mi mathvariant="script">L</mi><mi>i</mi></msub><mrow><mo fence="true">(</mo><mi>θ</mi><mo separator="true">,</mo><msub><mi mathvariant="script">D</mi><mi>i</mi></msub><mo fence="true">)</mo></mrow></mrow><annotation encoding="application/x-tex">\min _{\theta} \sum_{i=1}^{T} \mathscr{L}_{i}\left(\theta, \mathscr{D}_{i}\right)</annotation></semantics></math></span></p>
<p>The typical procedure:</p>
<ol>
<li>Sample a minibatch of tasks <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="script">B</mi><mo>∼</mo><mrow><mo fence="true">{</mo><msub><mi mathvariant="script">T</mi><mi>i</mi></msub><mo fence="true">}</mo></mrow></mrow><annotation encoding="application/x-tex">\mathscr{B} \sim\left\{\mathscr{T}_{i}\right\}</annotation></semantics></math></span></li>
<li>Sample a minibatch of data from each task <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msubsup><mi mathvariant="script">D</mi><mi>i</mi><mi>b</mi></msubsup><mo>∼</mo><msub><mi mathvariant="script">D</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">\mathscr{D}_{i}^{b} \sim \mathscr{D}_{i}</annotation></semantics></math></span></li>
<li>Compute the loss on each minibatch-task: <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi mathvariant="script">L</mi><mo>^</mo></mover><mo stretchy="false">(</mo><mi>θ</mi><mo separator="true">,</mo><mi mathvariant="script">B</mi><mo stretchy="false">)</mo><mo>=</mo><msub><mo>∑</mo><mrow><msub><mi mathvariant="script">T</mi><mi>k</mi></msub><mo>∈</mo><mi mathvariant="script">B</mi></mrow></msub><msub><mi mathvariant="script">L</mi><mi>k</mi></msub><mrow><mo fence="true">(</mo><mi>θ</mi><mo separator="true">,</mo><msubsup><mi mathvariant="script">D</mi><mi>k</mi><mi>b</mi></msubsup><mo fence="true">)</mo></mrow></mrow><annotation encoding="application/x-tex">\hat{\mathscr{L}}(\theta, \mathscr{B})=\sum_{\mathcal{T}_{k} \in \mathscr{B}} \mathscr{L}_{k}\left(\theta, \mathscr{D}_{k}^{b}\right)</annotation></semantics></math></span></li>
<li>Backpropagate to compute gradients <span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="normal">∇</mi><mi>θ</mi></msub><mover accent="true"><mi mathvariant="script">L</mi><mo>^</mo></mover></mrow><annotation encoding="application/x-tex">\nabla_{\theta} \hat{\mathscr{L}}</annotation></semantics></math></span></li>
<li>Update the gradients with your favorite optimizer</li>
</ol>
<p>Note: this ensures that tasks are sampled uniformly regardless of their data size.</p>
<p>Tip: for regression tasks, make sure task labels are on the same scale.</p>
<h3 id="challenge">Challenge</h3>
<ol>
<li>Negative transfer</li>
</ol>
<p>Multi-task training on CIFAR-100 performs worse than training tasks independently.</p>
<p>Possible causes:</p>
<ul>
<li>optimization challenges
<ul>
<li>interference between different tasks</li>
<li>different learning rates across tasks</li>
</ul>
</li>
<li>limited expressive capacity
<ul>
<li>multi-task networks are large</li>
</ul>
</li>
</ul>
<p>Solution:</p>
<p>share less across tasks (soft parameter sharing)</p>
<ul>
<li>allows for more fluid degrees of parameter sharing (advantage)</li>
<li>yet another set of design decisions/hyperparameters (drawback)</li>
</ul>
<span class="katex"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mrow><mi>min</mi><mo>⁡</mo></mrow><mrow><msup><mi>θ</mi><mrow><mi>s</mi><mi>h</mi></mrow></msup><mo separator="true">,</mo><msup><mi>θ</mi><mn>1</mn></msup><mo separator="true">,</mo><mo>…</mo><mo separator="true">,</mo><msup><mi>θ</mi><mi>T</mi></msup></mrow></msub><msubsup><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></msubsup><msub><mi mathvariant="script">L</mi><mi>i</mi></msub><mrow><mo fence="true">(</mo><mrow><mo fence="true">{</mo><msup><mi>θ</mi><mrow><mi>s</mi><mi>h</mi></mrow></msup><mo separator="true">,</mo><msup><mi>θ</mi><mi>i</mi></msup><mo fence="true">}</mo></mrow><mo separator="true">,</mo><msub><mi mathvariant="script">D</mi><mi>i</mi></msub><mo fence="true">)</mo></mrow><mo>+</mo><msubsup><mo>∑</mo><mrow><mi>t</mi><mo>=</mo><mn>1</mn></mrow><mi>T</mi></msubsup><mrow><mo fence="true">∥</mo><msup><mi>θ</mi><mi>t</mi></msup><mo>−</mo><msup><mi>θ</mi><msup><mi>t</mi><mo mathvariant="normal" lspace="0em" rspace="0em">′</mo></msup></msup><mo fence="true">∥</mo></mrow></mrow><annotation encoding="application/x-tex">
\min _{\theta^{sh}, \theta^{1}, \ldots, \theta^{T}} \sum_{i=1}^{T} \mathscr{L}_{i}\left(\left\{\theta^{s h}, \theta^{i}\right\}, \mathscr{D}_{i}\right)+\sum_{t=1}^{T}\left\|\theta^{t}-\theta^{t&#x27;}\right\|
</annotation></semantics></math></span><p>The latter term is soft parameter sharing: the difference between one task&rsquo;s parameters and the previous one is used as a regularization term, which effectively makes each task&rsquo;s parameters as similar as possible — i.e., the parameters are shared.</p>
<ol start="2">
<li>Overfitting</li>
</ol>
<p>Overfitting is usually caused by not sharing enough parameters; the solution is to share more. Intuitively, insufficient sharing makes each task overfit, which resembles independent training.</p>
<h2 id="meta-learning-basics">Meta-Learning Basics</h2>
<p>Two views of meta-learning:</p>
<ul>
<li>Mechanistic view
<ul>
<li>a deep neural network that can take in an entire dataset and make predictions on new data</li>
<li>the network is trained on a meta-dataset that contains different datasets for different tasks</li>
<li>this view makes it easy to implement a meta-learning algorithm</li>
</ul>
</li>
<li>Probabilistic view
<ul>
<li>extract prior knowledge from a series of meta-learning tasks</li>
<li>use a small amount of data plus prior information to infer a relatively effective posterior</li>
<li>this view leads to a better understanding of meta-learning algorithms</li>
</ul>
</li>
</ul>
<h3 id="problem-definitions">Problem definitions</h3>
<p>First, recall supervised learning:</p>
<p><img alt="cs330-1-7.png" loading="lazy" src="http://cdn.ereebay.me/hexo/cs330-1-7.png"></p>
<p>Existing issues:</p>
<ul>
<li>requires a large amount of labeled data</li>
<li>labels are very limited for some tasks nowadays</li>
</ul>
<p>To be continued</p>
]]></content:encoded></item></channel></rss>