GIGAスクール構想によって、学校で1人1台端末を使う環境は大きく広がりました。 文部科学省も、1人1台端末と高速大容量の通信環境を、 学習活動や教育データの活用につなげることを進めています。
ただ、英語の授業で端末を使うとき、 調べ学習やデジタル教科書の閲覧だけで終わってしまうこともあります。 私がスピーキング授業で端末を使うときに考えているのは、 「端末に何をさせるか」ではなく、「生徒が話す時間をどう増やすか」です。
1人1台端末の強みは、教師の代わりになることではありません。
同じ時間に多くの生徒が練習できること、
一度で終わらず再挑戦しやすいこと、
個人練習と対人活動をつなぎやすいことにあります。
1. ペアワークと端末練習は「どちらが上」ではない
元の記事では、ペアワークには限界があり、 端末に向かって話すことで「完全に心理的安全性の高い空間」が作れると書いていました。 今は、この説明は単純すぎると考えています。
ペアワークには、人の反応を見ながら言い換えたり、 聞き返したりするという重要な学習があります。 一方で、英語を人前で話すことに強い緊張を感じる生徒にとっては、 最初の一回を個人で練習できる方が参加しやすい場合があります。
だから私は、 個人練習で準備する → ペアで使う → 必要なら再度個人練習する という行き来を作るようにしています。 端末はペア活動を置き換える道具ではなく、 対話に入る前の練習機会を増やす道具として使えます。
2. 音声認識は「発音の先生」ではない
PicSpeakでは、Web Speech APIの音声認識を使って、 生徒の発話を文字として認識し、練習上のフィードバックを返します。 ただし、音声認識が拾った結果を、そのまま「発音の正しさ」と考えることはできません。
認識結果は、発音だけでなく、 マイク、周囲の騒音、通信環境、ブラウザ、端末、話す速度などの影響を受けます。 MDNでも、SpeechRecognitionはすべての主要ブラウザで同じように利用できる機能ではなく、 一部のブラウザではサーバー側の認識サービスを利用することが説明されています。
そのためPicSpeakの表示は、 標準化された発音テストや英語力測定ではなく、同じ課題を繰り返すための練習上の目安 として扱うのが適切です。
3. 私が授業で最初に確認するのは「マイクが反応しているか」
端末を使ったスピーキング活動では、 学習内容以前に「自分の声が認識されているか分からない」という問題が起きます。 生徒が一生懸命話しているのに画面が反応しなければ、 努力が結果に反映されず、活動を続ける意欲も下がります。
そこで私は、本番課題に入る前に、 短い一文や数語を話して、画面が反応するか確認する時間を取ることが重要だと考えています。 反応しなければ、次の課題へ進ませるのではなく、 マイク許可、ブラウザ、音量、周囲の騒音などを確認します。
- マイクの使用許可が出ているか確認する
- 最初に短いテスト発話を行い、認識反応を見る
- イヤホンマイクが必要な環境か判断する
- 反応しない場合の代替手段を教師が事前に用意する
- 認識されないことを「英語が下手」と生徒に受け取らせない
4. 45〜50分で組むなら、私はこう使う
PicSpeakを使うこと自体を授業の目的にするのではなく、 「個人で準備し、人と話し、もう一度試す」という流れの中に置きます。 たとえば、1枚の写真を使う授業なら次のように組めます。
0〜3分|接続・マイクチェック
“I can see a boy.” など短い発話で、端末が声を拾っているか確認します。
3〜8分|全体Warm-up
プロジェクターに同じ写真を出し、名詞・動詞・形容詞をクラスで短く出します。
8〜15分|1回目の個人発話
端末を使って、自分が今言える範囲で写真を説明します。最初から高得点を狙わせません。
15〜23分|Mini Input
教師が全員に役立ちそうな表現を3〜5個だけ取り上げます。生徒同士の良い表現を共有しても構いません。
23〜30分|2回目の個人発話
追加した表現を一つ以上使い、同じ写真でもう一度話します。前回との違いを自分で確認します。
30〜40分|ペアへの転用
端末を閉じ、別の写真や同じテーマを使って相手に説明します。必要なら聞き返しも入れます。
40〜45分|振り返り
「最初は言えなかったが、最後には言えた表現」を一つ書かせます。残り時間は教師のフィードバックや機器トラブルの調整に使います。
時間配分はクラスによって変わります。 大切なのは、端末を触っている時間を増やすことではなく、 1回目と2回目の発話の間に学習が起きる構造を作ることです。
5. 「全員が必ず話す」ではなく「全員に話す機会を作る」
元記事では、「全員が恥ずかしがらずに、必ず英語を話す」と書いていました。 しかし、どの活動でも全員が同じように話せるとは限りません。 音声認識が苦手な生徒、発話に強い不安がある生徒、 聴覚・発話面で個別の配慮が必要な生徒もいます。
教師側ができるのは、 全員の発話を保証することではなく、 できるだけ多くの生徒が参加できる複数の入口を用意することです。
- 端末に向かって一人で話す
- 先に英文を書いてから読む
- キーワードだけ見ながら話す
- ペアで交互に一文ずつ話す
- 教師や支援者と少人数で練習する
「端末を使えば心理的安全性が生まれる」のではなく、 生徒が自分に合った参加方法を選べる授業設計が、 安心して挑戦するための条件の一つになると考えています。
6. 教員の役割は「教える」から消えるのではない
ICTを使うと、「教師は教える人からファシリテーターへ変わる」と言われることがあります。 私も、生徒を観察し、励まし、活動を支える役割は重要だと思います。
しかし、教師が教える役割までなくなるわけではありません。 生徒が共通して困っている表現を短く教えること、 音声認識の結果を過大評価しないよう説明すること、 ペア活動でコミュニケーションが成立するよう支援すること。 これらは教師にしかできない判断です。
音声認識に「評価」を丸ごと任せるのではなく、 機械には反復練習を支えてもらい、 教師は生徒の変化やコミュニケーション全体を見る。 私はその役割分担の方が、学校現場では現実的だと考えています。
7. 端末導入で見るべき数字は「スコア」だけではない
端末を使った活動が有効だったかを見るとき、 PicSpeakの得点だけを見ても十分ではありません。 授業者としては、次のような点も確認したいところです。
- 1人あたり何回発話できたか
- 1回目より2回目で使える表現が増えたか
- 端末練習後のペア活動で発話が続いたか
- 機器トラブルで参加できなかった生徒はいなかったか
- 生徒自身が「前より言えるようになった点」を説明できるか
1人1台端末の価値は、画面の中にあるのではありません。 これまで待ち時間になっていた時間を、生徒が練習する時間へ変えられるかにあります。 PicSpeakも、そのための一つの選択肢として改善を続けています。
参考資料
Japan's GIGA School initiative has greatly expanded access to one-device-per-student learning environments. MEXT continues to promote the effective use of these devices and school networks for learning activities and educational improvement.
In English classes, however, devices can easily become tools mainly for searching, reading digital materials, or completing worksheets. When I use devices for speaking lessons, my main question is not, “What can the device do?” It is “How can this design increase the time students actually speak?”
The strength of one-to-one devices is not that they replace teachers.
They can allow many students to practise at the same time,
make retrying easier, and connect individual preparation with interpersonal communication.
1. Pair Work and Device Practice Are Complementary
The previous version of this article suggested that pair work had major limitations and that speaking to a device could create a completely psychologically safe space. I now think that was too simple.
Pair work offers something devices cannot: learners see another person's reactions, ask for clarification, and adjust their message in real time. On the other hand, some students may find it easier to make a first attempt privately before speaking to a classmate.
I therefore prefer a cycle such as individual preparation → pair interaction → individual retry when needed. The device does not replace interaction; it can help students prepare for it.
2. Speech Recognition Is Not a Pronunciation Teacher
PicSpeak uses the Web Speech API to recognize speech as text and provide practice feedback. But recognition results should not be treated as a direct measurement of pronunciation accuracy.
Results can be affected by microphones, background noise, connectivity, browsers, devices, and speaking speed as well as pronunciation. MDN also notes that SpeechRecognition has limited browser availability and that some browsers use server-based recognition services.
PicSpeak feedback should therefore be treated as a practice reference for repeated attempts, not a standardized pronunciation or proficiency test.
3. Before the Task, I Check Whether the Microphone Is Actually Responding
A practical problem comes before pedagogy: students may not know whether the system is hearing them. If a learner speaks seriously but the screen does not respond, the activity can quickly become frustrating.
For this reason, I think a short microphone check should come before the main task. Students can say one short sentence or a few words and confirm that the system responds before moving on.
- Confirm microphone permission.
- Run a short test utterance before the main activity.
- Decide whether headsets are needed in the classroom environment.
- Prepare an alternative route when recognition does not work.
- Do not let students interpret recognition failure as “bad English.”
4. A 45–50 Minute Lesson Example
I do not make PicSpeak the purpose of the lesson. I place it inside a sequence of preparing individually, communicating with another person, and trying again.
0–3 min | Connection & microphone check
Use a short sentence such as “I can see a boy” to confirm that speech is being detected.
3–8 min | Whole-class warm-up
Show one picture and quickly collect useful nouns, verbs, and adjectives.
8–15 min | First individual attempt
Students describe the image using what they can currently say. The goal is not a high score.
15–23 min | Mini input
The teacher introduces only three to five expressions that will help many students.
23–30 min | Second individual attempt
Students try the same image again and deliberately use at least one new expression.
30–40 min | Transfer to pair interaction
Close the devices and use the language with a partner, using the same topic or a new picture.
40–45 min | Reflection
Students record one expression they could not say at first but could use by the end.
The timing will vary by class. What matters is not increasing screen time, but creating a learning event between the first and second attempts.
5. From “Everyone Must Speak” to “Everyone Gets a Way to Participate”
The original article promised that every student would speak English without hesitation. No activity can guarantee that. Some students may experience strong speaking anxiety, have difficulty with speech recognition, or require individual accommodations.
The teacher's role is therefore not to guarantee identical participation, but to create multiple entry points into speaking practice.
- Speak privately to a device.
- Write a sentence before saying it aloud.
- Speak from keywords rather than a full script.
- Build a description one sentence at a time with a partner.
- Practise in a small group with teacher support.
6. Technology Does Not Remove the Teacher's Teaching Role
ICT can reduce some repetitive work and give students more practice opportunities, but teachers still make important instructional decisions.
We decide what language should be taught explicitly, explain the limitations of speech recognition, observe students who are struggling, and help pair interaction become meaningful communication.
Rather than delegating “assessment” entirely to a machine, I prefer to let technology support repetition while the teacher looks at the learner's broader development and communication.
7. The Important Numbers Are Not Only Scores
When evaluating a device-supported lesson, PicSpeak scores alone are not enough. I also want to know:
- How many speaking attempts did each learner get?
- Did the second attempt include more usable language than the first?
- Did pair interaction continue more easily after individual practice?
- Did technical problems prevent anyone from participating?
- Can students explain what they can now say that they could not say before?
The value of one-to-one devices is not simply on the screen. It lies in whether we can turn waiting time into meaningful practice time. PicSpeak is one option I continue developing for that purpose.